RuleTreeRank (RTR) is an interpretable Learning-to-Rank framework that models ranking as a two-stage process: it first groups items with a shallow rule tree, then refines the score locally through instance-based comparisons.
RTR is designed for query-supported ranking problems. In candidate screening, for example, a query is a job offer and each item is a candidate. The same candidate can be relevant for one job and irrelevant for another, so the final ranking is induced within each query by sorting the RTR scores in descending order.
- Stage I: rule-based grouping. A shallow
RuleTreeRegressorpartitions the feature space into interpretable leaves and assigns each item a coarse score. - Stage II: local refinement. Inside each leaf, RTR learns an interpretable pairwise distance model and uses a local k-NN correction over the Stage-I residuals.
- Explanations. The first stage gives a decision path, while the second stage exposes the local neighbours and pairwise rules used for the correction.
The package also exposes MixedRTR, a query-aware variant that trains the local refinement by both tree leaf and query when query-specific historical context is required.
pip install ruletreerankLatest development version:
pip install git+https://github.com/jacons/RuleTreeRank.gitFrom a clone, for development:
git clone https://github.com/jacons/RuleTreeRank.git
cd RuleTreeRank
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytestRequires Python 3.10 or newer.
import numpy as np
from RuleTree import RuleTreeRegressor
from ltr_utility import ModelParam
from ruletreerank import KNNRegFast, PairwiseDistanceTree, RuleTreeRank
X_train = np.random.default_rng(0).normal(size=(100, 5))
y_train = X_train @ np.array([1.4, -0.7, 0.5, 0.0, 0.2])
q_train = np.repeat(np.arange(10), 10)
model = RuleTreeRank(
distance_f=ModelParam(PairwiseDistanceTree, {
"base_regressor": ModelParam(RuleTreeRegressor, {"max_depth": 2, "random_state": 0}),
"feature_concat": False,
"feature_diff": True,
"feature_sq_diff": True,
"subsample": 0.5,
}),
aggregation_f=ModelParam(KNNRegFast, {"n_neighbors": 5, "n_jobs": 1}),
base_regressor=RuleTreeRegressor(max_depth=2, random_state=0),
dist_objective="residuals",
)
model.fit(X_train, y_train, q_train)
scores = model.predict(X_train, q=q_train, output="full")
ranking_for_query_0 = np.argsort(-scores[q_train == 0])Ranking data is represented as a collection of queries and their items:
RTR learns a two-stage pointwise scoring function:
The first term,
Training residuals are then computed as
The final ranking for a query is obtained by sorting items by
- examples/min_example.ipynb: minimal RTR example on a scikit-learn dataset, with tables and plots showing the double-stage behaviour.
- examples/synthetic_ranking_task.ipynb: synthetic ranking task with a hidden linear scoring function and relevance labels induced by ranking.
src/ruletreerank/: core RTR, MixedRTR, pairwise distance tree, and fast k-NN components.src/ltr_utility/: shared interfaces, dataset utilities, query splitting, clustering, and explanation helpers.examples/: executable notebooks for minimal and synthetic ranking workflows.tests/: smoke test covering the two-stage fit/predict path.
