Compositional variable selection in quantile regression for microbiome data with false discovery rate control

Advancement in high-throughput sequencing technologies has stimulated intensive research interests to identify specific microbial taxa that are associated with disease conditions. Such knowledge is invaluable both from the perspective of understanding biology and from the biomedical perspective of therapeutic development, as the microbiome is inherently modifiable. Despite availability of massive data, analysis of microbiome compositional data remains difficult. The nature that relative abundances of all components of a microbial community sum to one poses challenges for statistical analysis, especially in high-dimensional settings, where a common research theme is to select a small fraction of signals from amid many noisy features. Motivated by studies examining the role of microbiome in host transcriptomics, we propose a novel approach to identify microbial taxa that are associated with host gene expressions. Besides accommodating compositional nature of microbiome data, our method both achieves FDR-controlled variable selection, and captures heterogeneity due to either heteroscedastic variance or non-location-scale covariate effects displayed in the motivating dataset. We demonstrate the superior performance of our method using extensive numerical simulation studies and then apply it to real-world microbiome data analysis to gain novel biological insights that are missed by traditional mean-based linear regression analysis.

Files

Metadata

Work Title Compositional variable selection in quantile regression for microbiome data with false discovery rate control
Access
Open Access
Creators
  1. Runze Li
  2. Jin Mu
  3. Songshan Yang
  4. Cong Ye
  5. Xiang Zhan
Keyword
  1. compositional covariates
  2. false discovery rate control
  3. microbiome data analysis
  4. quantile regression
  5. SCAD penalty
  6. variable selection
License In Copyright (Rights Reserved)
Work Type Article
Publisher
  1. Statistical Analysis and Data Mining
Publication Date March 27, 2024
Publisher Identifier (DOI)
  1. https://doi.org/10.1002/sam.11674
Deposited October 07, 2024

Versions

Analytics

Collections

This resource is currently not in any collection.

Work History

Version 1
published

  • Created
  • Added Statistical_Analysis_-_2024_-_Li_-_Compositional_variable_selection_in_quantile_regression_for_microbiome_data_with_false-2.pdf
  • Added Creator Runze Li
  • Added Creator Jin Mu
  • Added Creator Songshan Yang
  • Added Creator Cong Ye
  • Added Creator Xiang Zhan
  • Published
  • Updated
  • Updated Keyword, Publication Date Show Changes
    Keyword
    • compositional covariates, false discovery rate control, microbiome data analysis, quantile regression, SCAD penalty, variable selection
    Publication Date
    • 2024-04-01
    • 2024-03-27