ProteoParc: A tool to generate protein reference databases for ancient and non-model organisms
Carrillo-Martin G, Krueger J, Marques-Bonet T, Lizano E.
Abstract
20 Over the last few years, the increasing interest in analysing the proteome of extinct 21 and non-model organisms has generated a new field of research expanding the scope 22 of proteomics. The lack of curated databases and/or molecular data from these 23 organisms forces researchers to manually search in different public repositories for 24 related protein sequences, either for MS/MS peptide identification or ZooMS marker 25 annotation. This can lead to format incongruences and hinder reproducibility between authors: Email:guillermo.carrillo@upf.edu (G.C.M.); esther.lizano@upf.edu (E.L.); tomas.marques@upf.edu (T.M.B) bioRxiv preprint doi: https://doi.org/10.1101/2025.07.31.667843; this version posted August 2, 2025. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-NC-ND 4.0 International license. 26 studies. To address this issue, we introduce ProteoParc, a user-friendly software that 27 generates reference databases by systematically downloading and processing protein 28 sequences from the most widely used public repositories. The pipeline’s output is a 29 non-redundant protein database, formatted to be interpreted by typical peptide 30 identification software. Moreover, the user can adjust the database dimension and 31 composition by applying different criteria to include only a certain number of genes or 32 species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful 33 for future paleoproteomics analysis in ancient samples related to understudied 34 organisms. 35
This page indexes the study's public bibliographic record. The full text belongs to the journal; follow the DOI above to read it at the source.