In metabolomics data, like other -omics data, normalization is an important part of the data processing. The goal of normalization is to reduce the variation from non-biological sources (such as instrument batch effects), while maintaining the biological variation. Many normalization techniques make adjustments to each sample. One common method is to adjust each sample by its Total Ion Current (TIC), i.e. for each feature in the sample, divide its intensity value by the total for the sample. Because many of the assumptions of these methods are dubious in metabolomics data sets, we compare these methods to two methods that make adjustments separately for each metabolite, rather than for each sample. These two methods are the following: 1) for each metabolite, divide its value by the median level in bridge samples (BRDG); 2) for each metabolite divide its value by the median across the experimental samples (MED). These methods were assessed by comparing the correlation of the normalized values to the values from targeted assays for a subset of metabolites in a large human plasma data set. The BRDG and MED normalization techniques greatly outperformed the other methods, which often performed worse than performing no normalization at all.
Deininger, S.O., et al. (2011) Normalization in MALDI-TOF Imaging Datasets of Proteins: Practical Considerations. Analytical and Bioanalytical Chemistry, 401, 167-181. https://doi.org/10.1007/s00216-011-4929-z
Warrack, B.M., et al. (2009) Normalization Strategies for Metabonomic Analysis of Urine Samples. Journal of Chromatography B-Analytical Technologies in the Biomedical and Life Sciences, 877, 547-552. https://doi.org/10.1016/j.jchromb.2009.01.007
Webb-Robertson, B.J., Matzke, M.M., Jacobs, J.M., Pounds, J.G. and Waters, K.M. (2011) A Statistical Selection Strategy for Normalization Procedures in LC-MS Proteomics Experiments through Dataset-Dependent Ranking of Normalization Scaling Factors. Proteomics, 11, 4736-4741. https://doi.org/10.1002/pmic.201100078
Dieterle, F., Ross, A., Schlotterbeck, G. and Senn, H. (2006) Probabilistic Quotient Normalization as Robust Method to Account for Dilution of Complex Biological Mixtures. Application in 1H NMR Metabonomics. Analytical Chemistry, 78, 4281-4290. https://doi.org/10.1021/ac051632c
Yang, Y.H., et al. (2002) Normalization for cDNA Microarray Data: A Robust Composite Method Addressing Single and Multiple Slide Systematic Variation. Nucleic Acids Research, 30, e15. https://doi.org/10.1093/nar/30.4.e15
Sysi-Aho, M., Katajamaa, M., Yetukuri, L. and Oresic, M. (2007) Normalization Method for Metabolomics Data Using Optimal Selection of Multiple Internal Standards. BMC Bioinformatics, 8, 93. https://doi.org/10.1186/1471-2105-8-93
Nezami Ranjbar, M.R., Zhao, Y., Tadesse, M.G., Wang, Y. and Ressom, H.W. (2013) Gaussian Process Regression Model for Normalization of LC-MS Data Using Scan-Level Information. Proteome Science, 11, S13. https://doi.org/10.1186/1477-5956-11-S1-S13
Altman, D.G. and Bland, J.M. (1983) Measurement in Medicine: The Analysis of Method Comparison Studies. The Statistician, 32, 307-317. https://doi.org/10.2307/2987937
Bland, J.M. and Altman, D.G. (1986) Statistical Methods for Assessing Agreement between Two Methods of Clinical Measurement. The Lancet, 1, 307-310. https://doi.org/10.1016/S0140-6736(86)90837-8
Astrand, M. (2003) Contrast Normalization of Oligonucleotide Arrays. Journal of Computational Biology, 10, 95-102. https://doi.org/10.1089/106652703763255697
Bolstad, B.M., Irazarry, R.A., Astrand, M.T. and Speed, P. (2003) A Comparison of Normalization Methods for High Density Oligonucleotide Array Data Based on Variance and Bias. Bioinformatics, 19, 185-193. https://doi.org/10.1093/bioinformatics/19.2.185
Rocke, D.M. and Lorenzato, S. (1995) A Two-Component Model for Measurement Error in Analytical Chemistry. Technometrics, 37, 176-184. https://doi.org/10.1080/00401706.1995.10484302
Peters, F.T. and Maurer, H.H. (2007) Systematic Comparison of Bias and Precision Data Obtained with Multiple-Point and One-Point Calibration in Six Validated Multianalyte Assays for Quantification of Drugs in Human Plasma. Analytical Chemistry, 79, 4967-4976. https://doi.org/10.1021/ac070054s
Kenny, L.C., et al. (2010) Robust Early Pregnancy Prediction of Later Preeclampsia Using Metabolomic Biomarkers. Hypertension, 56, 741-749. https://doi.org/10.1161/HYPERTENSIONAHA.110.157297
Zhou, B., Xiao, J.F., Tuli, L. and Ressom, H.W. (2012) LC-MS-Based Metabolomics. Molecular BioSystems, 8, 470-481. https://doi.org/10.1039/C1MB05350G
Henkin, et al. (2003) Genetic Epidemiology of Insulin Resistance and Visceral Adiposity. The IRAS Family Study Design and Methods. Annals of Epidemiology, 13, 211-217. https://doi.org/10.1016/S1047-2797(02)00412-X
Long, T., et al. (2017) Whole-Genome Sequencing Identifies Common-to-Rare Variants Associated with Human Blood Metabolites. Nature Genetics, 49, 568-578. https://doi.org/10.1038/ng.3809
Cobb, J., et al. (2015) A Novel Test for IGT Utilizing Metabolite Markers of Glucose Tolerance. Journal of Diabetes Science and Technology, 9, 69-76. https://doi.org/10.1177/1932296814553622
Kaddurah-Daouk, R., et al. (2011) Enteric Microbiome Metabolites Correlate with Response to Simvastatin Treatment. PLoS ONE, 6, e25482. https://doi.org/10.1371/journal.pone.0025482
R Core Team (2017) R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna. https://www.R-project.org/
Ritchie, M.E., et al. (2015) Limma Powers Differential Expression Analyses for RNA-Sequencing and Microarray Studies. Nucleic Acids Research, 43, e47. https://doi.org/10.1093/nar/gkv007
Hochrein, J., et al. (2015) Data Normalization of (1)H NMR Metabolite Fingerprinting Data Sets in the Presence of Unbalanced Metabolite Regulation. Journal of Proteome Research, 14, 3217-3228. https://doi.org/10.1021/acs.jproteome.5b00192