Finding phylogeny-aware and biologically meaningful averages of metagenomic samples: L2UniFrac

Motivation Metagenomic samples have high spatiotemporal variability. Hence, it is useful to summarize and characterize the microbial makeup of a given environment in a way that is biologically reasonable and interpretable. The UniFrac metric has been a robust and widely used metric for measuring the variability between metagenomic samples. We propose that the characterization of metagenomic environments can be improved by finding the average, a.k.a. the barycenter, among the samples with respect to the UniFrac distance. However, it is possible that such a UniFrac-average includes negative entries, making it no longer a valid representation of a metagenomic community.

Results To overcome this intrinsic issue, we propose a special version of the UniFrac metric, termed L2UniFrac, which inherits the phylogenetic nature of the traditional UniFrac and with respect to which one can easily compute the average, producing biologically meaningful environment-specific “representative samples.” We demonstrate the usefulness of such representative samples as well as the extended usage of L2UniFrac in efficient clustering of metagenomic samples, and provide mathematical characterizations and proofs to the desired properties of L2UniFrac.

Availability and implementation A prototype implementation is provided at https://github.com/KoslickiLab/L2-UniFrac.git. All figures, data, and analysis can be reproduced at https://github.com/KoslickiLab/L2-UniFrac-Paper

Files

Metadata

Work Title Finding phylogeny-aware and biologically meaningful averages of metagenomic samples: L2UniFrac
Access
Open Access
Creators
  1. Wei Wei
  2. Andrew Millward
  3. David Koslicki
License In Copyright (Rights Reserved)
Work Type Article
Publisher
  1. Bioinformatics
Publication Date June 30, 2023
Publisher Identifier (DOI)
  1. https://doi.org/10.1093/bioinformatics/btad238
Deposited March 22, 2024

Versions

Analytics

Collections

This resource is currently not in any collection.

Work History

Version 1
published

  • Created
  • Added btad238.pdf
  • Added Creator Wei Wei
  • Added Creator Andrew Millward
  • Added Creator David Koslicki
  • Published
  • Updated
  • Updated Publisher, Publication Date Show Changes
    Publisher
    • ISMB 2023
    • Bioinformatics
    Publication Date
    • 2023-06-01
    • 2023-06-30
  • Updated Work Title Show Changes
    Work Title
    • L2-UniFrac: taking averages with respect to a phylogenetically informed metric
    • Finding phylogeny-aware and biologically meaningful averages of metagenomic samples: L2UniFrac