I will do my best to maintain these. You can let me know if any of the links is dead.
Phlag
Şapcı, A. O. B., Arasti, S., Braun, E. L., & Mirarab, S. (2026). “Phlag: Scalable detection of genomics regions with unexplained phylogenetic heterogeneity.” Bioinformatics, 42 (Supplement 1), btag273. https://doi.org/10.1093/bioinformatics/btag273
- Gene trees simulated using msprime and simulation experiments: phlag-avian-simulations
- Analysis conducted on the mammalian phylogeny (Foley et al., 2023): phlag-mammalian-analysis
- Experiments on the Stiller et al., 2024 avian phylogeny: phlag-avian-analysis
- Phlag benchmarking results, together with the simulated ARGs (experiments E1 and E2), and the supporting data for the avian dataset (experiment E3) and the mammalian dataset (experiment E4): data-repository
krepp
Şapcı, Ali Osman Berk, and Siavash Mirarab. 2026. “krepp: a k-mer-based maximum pseudo-likelihood method for estimating read distances and genome-wide phylogenetic placement.” Genome Biology 27 (1): 108. https://doi.org/10.1186/s13059-026-03999-y.
- The software and its documentation: GitHub and Bioconda
- Reference indexes: Wiki
- Data sets for distance estimation and placement experiments: Dryad
- Results and scripts used in analysis: GitHub
KRANK
Şapcı, A.O.B. and Mirarab, S. (2024) “Memory-bound k-mer selection for large and evolutionary diverse reference libraries”, Genome Research, p. gr.279339.124. Available at: https://doi.org/10.1101/gr.279339.124.
- All data is available on Dryad.
- A catalog of reference libraries is available at https://ter-trees.ucsd.edu/data/krank.
- Simulated reads used for read classification experiments can be found here.
- Smaller files (including results, and misc data) are available on GitHub.
- All releases of the software and its documentation can be found on GitHub.
CONSULT-II
Şapcı, A.O.B., Rachtman, E. and Mirarab, S. (2024) “CONSULT-II: Accurate taxonomic identification and profiling using locality-sensitive hashing”, Bioinformatics, p. btae150. Available at: https://doi.org/10.1093/bioinformatics/btae150.
- CONSULT-II reference libraries: 140Gb (high-sensitivity), 32Gb (robust and lightweight), 18Gb (lightweight).
- Simulated reads for classification experiments: reads from bacterial genomes and reads from archaeal genomes.
- Smaller files (including results, scripts and taxonomy files) are available on GitHub.
- All releases of the software and its documentation can be found on GitHub.