dataset · Zenodo (CERN European Organization for Nuclear Research)
A curated protein sequence dataset and associated reproducibility files have been assembled to support the integrated analysis of carbapenemases. The initial collection of 4,178 protein accession records was dereplicated into 3,349 exact unique amino acid sequences. Multi-method validation using CARD RGI, NCBI AMRFinderPlus, Pfam domains, catalytic motifs, and phylogenetics confirmed 3,246 unique proteins across 4,069 accessions. From this, a primary quality-controlled cohort of 2,548 full-length unique proteins was established alongside a secondary cohort of 698 unique proteins. The deposited resource provides sequence mappings, cluster information, taxonomic and temporal records, phylogenetic trees, and analysis scripts, with all source entries remaining fully traceable to UniProtKB and UniParc identifiers.
Carbapenemase enzymes cause resistance to critical antibiotics, making accurate identification essential. Public databases often contain duplicate entries or inconsistent annotations. By systematically validating sequences and establishing a high-confidence, non-redundant reference cohort, this resource helps researchers rely on standardised, reproducible genomic and protein data when studying antimicrobial resistance mechanisms.
The abstract does not indicate a commercial application pathway, presenting instead a curated reference dataset and reproducibility files for academic and bioinformatics research.
AI-generated from the published abstract. Always read the original work before citing.
This dataset contains curated protein-sequence data and reproducibility materials supporting an integrated analysis of carbapenemase annotation confidence, exact-sequence redundancy, phylogenetic family/lineage resolution, taxonomic representation, and temporal database representation. The source collection comprised 4,178 protein accession records, dereplicated into 3,349 exact unique amino-acid sequences. Integrated validation using CARD RGI, NCBI AMRFinderPlus, Pfam-domain evidence, class-specific catalytic motifs, and phylogenetic analyses retained 3,246 unique proteins representing 4,069 accession records. The primary analytical cohort comprised 2,548 full-length QC-passing unique proteins representing 3,294 accession records, while the supported secondary cohort comprised 698 unique proteins representing 775 accession records. The deposited archive includes accession-to-sequence mappings, exact-sequence clusters, cohort assignments, integrated annotation and validation outputs, final family/lineage assignments, phylogenetic alignments and trees, taxonomic and temporal analysis tables, Supplementary Tables S1–S20, analysis scripts, and reproducibility records. Source protein records remain traceable to their public UniProtKB and UniParc accessions. Taxonomic counts describe representation within the curated dataset and should not be interpreted as prevalence, transmission frequency, or host specificity. Temporal variables correspond to database-record creation/deposition dates and should not be interpreted as isolate-collection dates, biological emergence, or discovery dates.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.21882475
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.