Abstract
Raw and processed data from metabolomics and lipidomics experiments can be deposited in public repositories such as MetaboLights, MassIVE or MetabolomicsWorkbench. In addition, various public metabolite annotation resources including spectral libraries exist, albeit without standardized metadata or a common format.
The Spectra Bioconductor package provides a flexible infrastructure to handle and process mass spectrometry (MS) data from proteomics or metabolomics experiments. The strict separation of functionality for MS data handling, storage and representation from user-faced functions for data analysis facilitates expansion of Spectra to additional file types, data formats and resources. Dedicated backend implementations enable a direct access to MS data from public repositories simplifying hence integration of such data into reproducible data analysis workflows. Currently, the MsBackendMetaboLights Bioconductor package enables access to MS data from MetaboLights and analogous packages for data from MassIVE and MetabolomicsWorkbench are being developed as part of the project “Data analysis infrastructure MetaRbolomics4Galaxy”. Similarly, the MsBackendMassbank package adds support and enables access to the small molecule annotation resource MassBank, and, combined with the CompoundDb Bioconductor package, allows creation of small, redistributable SQLite annotation databases. Support for additional annotation resources, provided through Figshare or Zenodo, or formats, such as mzSpecLib, will be added in future.
Funding information: this work is co-funded by the Autonomous Province of Bolzano under the Joint Project MetaRBolomics4Galaxy (CUP: D53C25001030003).