Omics Codeathon General Application - October 2025
Organized by the African Society for Bioinformatics and Computational Biology (ASBCB) with support from the NIH Office of Data Science Strategy.
Single-cell chromatin accessibility sequencing (scATAC-seq) enables genome-wide profiling of regulatory elements at single-cell resolution. Traditional pipelines identify accessible regions first, then infer TF activity, limiting comprehensive understanding of regulatory programs driving cellular identity. This study develops a robust TF-centric machine learning framework to classify PBMC single-cell datasets using inferred chromVAR transcription factor activities. Our approach addresses data quality challenges through unsupervised redundancy filtering and class-imbalance handling via weighted loss functions, and employs multiple machine learning models for rigorous classification. The resulting computational pipeline enhances single-cell analysis capabilities and provides a systematic approach for discovering TF regulatory networks in immune cell populations.
The cellitac pipeline is designed for easy and direct use through official package managers, ensuring a reproducible environment for single-cell analysis.
- PyPI (Python): cellitac on PyPI
Figure 1. Workflow of the methods employed in this study
All scripts for the cellitac project (Python & R) are available in the repository:
👉 Browse the scripts: Scripts Running
- Public dataset: PBMC from a Healthy Donor (10k, 10x Genomics)
Cell types retained :
- B cells
- CD4+ T cells
- CD8+ T cells
- Dendritic cells
- Monocytes
- NK cells
- T cells
Final dataset after filtering: A total of 10,989 cells and 578 TF motifs distributed across 7 cell types.
| Pipeline Stage | Hardware Specification |
|---|---|
| Full Pipeline (Preprocessing & ML) | Personal Laptop: Intel Core 7 240H (16 Cores), 16 GB RAM, 1 TB SSD, NVIDIA GeForce RTX 5060 (8 GB VRAM) (WSL / Linux) |
all package versions (R - Python) specified for this project
To report an issue please use the issues page (https://github.com/omicscodeathon/cellitac/issues). Please check existing issues before submitting a new one.
You can offer to help with the further development of this project by making pull requests on this repo. To do so, fork this repository and make the proposed changes. Once completed and tested, submit a pull request to this repo.
| Name | Affiliation | Role |
|---|---|---|
| Rana Hamed | Student, School of Computing and Data Science, Badya University, Cairo, Egypt | Team Lead – Project Management |
| Syrus Semawule | African Center of Excellence in Bioinformatics and Data Intensive Sciences, The Infectious Disease Institute, Makerere University, Kampala, Uganda | Bioinformatician – Data Processing & Biological Annotation |
| Emmanuel Aroma | Department of Immunology and Molecular Biology, School of Biomedical Sciences, Makerere University, Kampala, Uganda | Bioinformatician – ML Modeling & Pipeline Control |
| Toheeb Jumah | Department of Human Anatomy, Faculty of Basic Medical Sciences, College of Medical Sciences, Ahmadu Bello University, Zaria, Nigeria | Bioinformatician – Manuscript Writing & ML Modeling |
| Olaitan I. Awe | African Society for Bioinformatics and Computational Biology (ASBCB), Cape Town, South Africa | Project Advisor |
📧 Rana Hamed Abu-Zeid : ranahamed2111@gmail.com
📧 Syrus Semawule : semawulesyrus@gmail.com
📧 Emmanuel Aroma : emmatitusaroma@gmail.com
📧 Toheeb Jumah : jumahtoheeb@gmail.com
📧 Olaitan I. Awe, Ph.D. : laitanawe@gmail.com
We thank the NIH Office of Data Science Strategy for their support before and during the October 2025 Omics Codeathon, co-organized with the African Society for Bioinformatics and Computational Biology (ASBCB).
We also thank Dr. Awe for his ongoing guidance and all collaborators who contributed to this project.
This project reflects a collaborative effort towards advancing integrative bioinformatics methods, and we look forward to its continued development and impact within the scientific community.



