Project overview

Aim

Create a custom reference protein database derived from microbiome sequencing data and use it to analyze differences in protein composition and functional profiles between groups of interest.

Key insights
  • For functional gut microbiome analysis, shotgun metagenomic sequencing provides the most suitable basis for custom protein database construction due to its species-level taxonomic resolution.
  • Although custom protein databases can also be generated from 16S rRNA sequencing data, this approach requires multiple analytical decisions i.a. regarding taxonomic assignment of ASVs/OTUs and selection of representative species-level matches.
  • Construction of a study-specific protein database is important because using overly broad reference databases may increase false positive protein identifications.
  • In gut microbiome proteomics, the reference protein database should ideally be based on complementary analyses, such as sequencing data from the same samples, because microbial composition and predicted proteomes vary substantially across individuals, age groups, intestinal tract regions, and other biological or environmental conditions.
  • Appropriate normalization is a critical step in proteomics analysis, as it enables reliable comparison of protein abundances between experimental groups.
  • Hypothesis-driven enrichment analysis facilitates biologically interpretable conclusions by focusing on specific pathways, enzymes, or functional categories, while broader exploratory analyses remain valuable for novel hypothesis generation.