Attribute Rank Enrichment Analysis
The goal of AREA is to link boolean attributes to rankable attributes via samples. In this way, we may be able to find rankable attributes that cause the boolean attributes or vice versa. Similar to correlation analysis, causality is not clear from this analysis alone. Instead, these linkages should be considered potentially causal, and downstream experiments should be used to determine causality.
In our example, we analyze gene expression data from individuals. These individuals have Down syndrome, and the boolean attributes are their associated medical conditions. Our objective is to identify genes that, when highly expressed, may influence the likelihood of specific medical conditions.
NES stands for Normalized Enrichment Score. A negative NES indicates that low levels of the rankable attribute (in our case gene expression) are associated with the true condition of the boolean attribute. Conversely, a positive NES suggests that high levels of the rankable attribute are associated with the true condition of the boolean attribute.
AREA requires Python 3.9 or later.
If you are on a shared supercomputer with SLURM, create a dedicated virtual environment for AREA:
module load python/3.9.15
mkdir -p ~/VENVs
cd ~/VENVs
python -m venv areavenv39
source areavenv39/bin/activate
pip install -r ~/AREA/requirements.txtActivate this environment before each AREA run:
source ~/VENVs/areavenv39/bin/activatecd AREA
pip install -e .For GPU acceleration via CuPy:
pip install -e ".[gpu]"AREA takes two CSV files as input. Both CSVs share the same samples (in our example, individuals), but one has boolean attributes for each sample and the other has rankable attributes. The output is a CSV with the statistical significance of the linkages between the boolean columns and the rank columns.
Both input CSV files must have a common sample column. In our case, that common sample name column is "Participant".
One of the CSV files needs to contain boolean attributes. In our case, the boolean attributes are the disease or disorder associated with each patient, which we call a comorbidity.
The other file must have columns that can be ranked by the values within the column. In our example, the rank file has genes and the expression level of those genes in each patient.
The output file lists each boolean attribute and rank column in pairs. The rest of the row has the scores for that pair, including NES, raw p-value, and four adjusted p-value columns (Bonferroni, Holm, Benjamini-Hochberg, and Benjamini-Yekutieli).
The items in the output file have NOT been filtered for significance. To filter for significance, pick an adjusted p-value column (four are provided) and apply your chosen cutoff.
There are three ways to run AREA after installation.
No install needed — run from inside the AREA/ directory:
python run_area.py --config area_config.yamlAfter pip install -e ., the area command is available anywhere:
area --config area_config.yamlpython -m area --config area_config.yamlCopy area_config.yaml and fill in your paths:
boolean_file: /path/to/bools.csv
rank_file: /path/to/ranks.csv
join_column: Participant
out_dir: /path/to/results/
threads: 4
gpu: false
verbose: falseThen run:
python run_area.py --config my_config.yamlAll parameters can also be passed directly on the command line:
python run_area.py \
-bf /path/to/bools.csv \
-rf /path/to/ranks.csv \
-jc Participant \
-od /path/to/results/Command-line flags always override values from the config file. This lets you set standard parameters in your config and override specific ones per run:
python run_area.py --config my_config.yaml --verbose --threads 8| Flag | Config key | Required | Description |
|---|---|---|---|
-bf / --boolean-file |
boolean_file |
Yes | CSV of binary attributes (samples x attributes) |
-rf / --rank-file |
rank_file |
Yes | CSV of continuous values (samples x features) |
-jc / --join-column |
join_column |
Yes | Column name shared between both input files |
-od / --out-dir |
out_dir |
Yes | Directory for all output files |
-t / --threads |
threads |
No | Number of parallel threads (default: 4) |
--gpu |
gpu |
No | Use GPU acceleration via CuPy (default: false) |
--keep-rank-columns |
keep_rank_columns |
No | Text file listing rank columns to include |
--keep-bool-columns |
keep_bool_columns |
No | Text file listing bool columns to include |
--keep-samples |
keep_samples |
No | Text file listing sample IDs to include |
--verbose |
verbose |
No | Print diagnostic output (default: false) |
-c / --config |
— | No | Path to a YAML config file |
Often a user will not want to run all data through AREA. If a boolean attribute is true for all samples or no samples, it will not produce meaningful results. Similarly, genes with no expression variance across samples are uninformative. AREA automatically excludes constant columns, but you can further limit which columns are analyzed using keep files:
--keep-bool-columns— a text file with one boolean column name per line--keep-rank-columns— a text file with one rank column name per line--keep-samples— a text file with one sample ID per line
AREA includes a test suite in tests/test_area.py that covers both unit tests and integration tests.
Place your test files in the testdata/ directory:
testdata/genes.csv— a rank file (samples x genes)testdata/comorbid_file.csv— a boolean file (samples x comorbidities)
Unit tests for the backend, config loader, enrichment math, and CLI parsing run without test data. The integration tests (planning, runner, correction, and the full end-to-end pipeline) require the test data files and will skip gracefully if they are not present.
From the AREA/ directory:
python run_tests.py
python run_tests.py -v # verbose output| Test class | Module tested | Requires test data |
|---|---|---|
TestBackend |
backend.py — CPU backend, to_numpy, trapz helper |
No |
TestConfig |
config.py — YAML loading, unknown-key rejection |
No |
TestEnrichment |
enrichment.py — enrichment score sign, permutation count, NES p-value range |
No |
TestCLI |
cli.py — argument parsing, missing-required errors, config/CLI override precedence |
No |
TestPlanning |
planning.py — plan file creation, row structure, constant-column exclusion |
Yes |
TestRunner |
runner.py — raw p-values output, column validation, p-values in [0, 1] |
Yes |
TestCorrection |
correction.py — adjusted file creation, Bonferroni/Holm/BH/BY columns, sort order |
Yes |
TestEndToEnd |
Full pipeline via cli.main() with --verbose and 2 threads |
Yes |
AREA/
├── pyproject.toml # Package metadata and dependencies
├── run_area.py # Top-level entry script
├── run_tests.py # Top-level test runner
├── area_config.yaml # Documented config template
├── examples/ # Example SLURM shell scripts
├── testdata/ # Test CSV files (not in repo)
│ ├── genes.csv
│ └── comorbid_file.csv
├── tests/ # Test suite
│ └── test_area.py
└── src/
└── area/
├── __init__.py # Package docstring
├── __main__.py # Enables "python -m area"
├── backend.py # GPU/CPU array backend selection
├── enrichment.py # Core math (enrichment score, permutations, NES)
├── planning.py # Run plan construction and column filtering
├── runner.py # Thread-pool orchestration
├── correction.py # Multiple-testing p-value correction
├── config.py # YAML config file loader
└── cli.py # Argument parsing and main() entry point
To discover what patterns AREA is best at finding, we created simulated data similar to our real data in notebook_examples/simulateddata-bothdirs.ipynb. This code creates a certain number of simulated genes based on the expression of real genes. We then create comorbidities that are biased in a sub-group of people with higher or lower expression.
The data used in our example code can be sourced from the INCLUDE Data Hub at https://portal.includedcc.org/ and anyone who would like to use the original data can register for an account on the data hub. Data used from the Human Trisomy Project was converted from kallisto to DESeq2 normalized counts using tximport.


