An opinionated framework for B2B revenue attribution. Pluggable rules, persistent learning, bring-your-own-CRM.
pip install cascade-attributionMost attribution stacks fail the same way: a vendor draws a straight line from first touch to closed deal and calls it a model. That line is fiction the moment a deal sees more than one touchpoint, which is every deal.
This framework takes a different bet. Attribution is a judgment problem with structure — a sequence of composable rules, each one answering a specific question about a specific deal, each one capable of being taught something new by the person actually closing the business.
The rules are Python functions. The corrections become persistent bias. The history is queryable. Swap your CRM adapter, keep the methodology.
I've been doing ops for nearly two decades and I've watched the same conversation happen in every company:
Leadership: "What's working?" Marketing: "Our tools say paid search." Sales: "No way. That deal came from the dinner we hosted in Q2." Data: "Actually it's a multi-touch journey. Here's a 14-tab spreadsheet."
Everyone is partially right. The dinner, the paid search click, the SDR follow-up, the customer referral, the podcast mention — they all contributed. But the acquisition event — the thing you'd credit if you had to name one — is knowable. Not by one vendor's model, but by a rep who watched it happen plus a framework that can learn from the rep's judgment.
That's what this is.
cascade-attribution/
├── README.md this
├── docs/
│ ├── methodology.md the framework, in detail
│ └── pattern-library.md 40 attribution edge cases with reasoning
├── cascade_attribution/
│ ├── rules/ the rule library — each rule is a Python function
│ ├── adapters/ Salesforce, HubSpot, CSV — pick one or write your own
│ ├── learning.py the correction loop
│ └── pipeline.py run rules → produce attribution → record outcome
├── fixtures/ fake deals + known-good attributions for tests
└── tests/
from cascade_attribution import attribute, load_rules
from cascade_attribution.adapters import CSVAdapter
# Your CRM, your rules
crm = CSVAdapter("opportunities.csv", "contacts.csv", "campaigns.csv")
rules = load_rules() # all 20+ default rules, or pass your own list
# Attribute one opportunity
opp = crm.get_opportunity("006...ABC")
result = attribute(opp, rules=rules, crm=crm)
print(result.campaign) # "Co-Sponsorship Invite"
print(result.lead_source) # "Referral - Customer"
print(result.touch_source) # None (invite path doesn't use tracking)
print(result.matched_rule) # "lo-invited-by-rea"
print(result.confidence) # 0.92
print(result.reasoning) # "Campaign membership present for Q2 invite campaign, campaign type 'LO Referrals' matched"When a rep disagrees with the attribution, you feed that back:
from cascade_attribution.learning import correct
correct(
opportunity_id="006...ABC",
actual_attribution={"campaign": "Group Demos - Q2", "lead_source": "Event"},
reason="attended the April group demo; invite campaign was ancillary",
)Next time a similar deal shows up, the rules library has a new bias. The correction log is queryable, versioned, and can be promoted to a permanent rule once it's fired enough times.
- Not a multi-touch attribution model. Multi-touch reports are fine for marketing spend allocation but terrible for deal-level judgment. This framework answers the question "what gets credit for this specific opp" — a different, more interesting problem.
- Not a CRM. Bring your own Salesforce, HubSpot, or just a CSV of opportunities. The framework is CRM-agnostic.
- Not prescriptive about your lead-source taxonomy. The reference rules ship with a reasonable default (LSD/TSD, Direct / Referral / Event / Nurture / Outbound / Paid / Organic), but every rule is a function — override the ones that don't fit, keep the ones that do.
- Not a silver bullet. This is a tool for people who already know ops work is judgment. If you want a vendor to tell you a number with high confidence and zero accountability, there are plenty of those.
- Run — a pipeline fires all rules in priority order against a deal.
- Match — the first rule that returns a non-null attribution wins.
- Record — the attribution + matched rule + reasoning gets logged.
- Correct — when a human disagrees, the correction feeds back as a persistent bias on future matches.
- Promote — when a correction fires N times consistently, it becomes a new rule in the library.
This is the same shape as every good classifier — rules + training data + a human in the loop. The difference is that the training data here is your reps telling you what actually happened, and the rules are legible Python functions that a director of ops can read on a Friday afternoon.
| Adapter | Status | What you need |
|---|---|---|
CSVAdapter |
✅ shipped | Three CSVs: opportunities, contacts, campaigns |
SalesforceAdapter |
📋 interface defined | simple-salesforce package + OAuth creds |
HubSpotAdapter |
📋 interface defined | hubspot-api-client package + API key |
The Adapter interface is small — five methods. Writing your own for
another CRM is an afternoon.
The default rule pack is a distillation of 20+ corrections I've watched actually happen at scale. Every rule is a pure function that takes a deal
- context and returns either
None(not my problem) or anAttribution(here's what fired).
Highlights:
co-sponsorship-invite— campaign typeLO Referralsor campaign name containsinvited by→ Referral / peer-to-peer path beats any ambient signalchurn-winback— outbound touch on a formerly-churned account with a reactivation campaign → Winback, even without a coupondirect-organic-mapping—src=direct+med=organic→ map toDirectlead-source (not Organic — common mistake)product-invite-vs-peer-referral— distinguishes between a Homebot-style product invite and a peer-to-peer referral by inspecting the source-data token, not just the channelnurture-driven-signup— if a contact had nurture touches before filling the signup form, the nurture wins — the form is fulfillment, not acquisition
See docs/pattern-library.md for the full list, each with a worked
example and the reasoning behind it.
- Rules are functions, not config. I tried the "YAML-based rules
engine" thing. It always collapses when you hit the rule that needs an
ifinside anif. Python functions keep everything inspectable and testable. - Corrections are first-class. Every attribution gets recorded. Every correction gets logged with a human-readable reason. The log is the actual asset — more valuable than any single attribution.
- No ML. On purpose. A logistic regression predicting "probability this deal came from paid search" is worse than a deterministic rule a rep can override. Judgment > inference at this data scale.
pip install cascade-attribution
# or from source
git clone https://github.com/swsounds42/cascade-attribution.git
cd cascade-attribution
pip install -e .Python 3.10+.
- docs/methodology.md — the framework, in detail
- docs/pattern-library.md — 40 edge cases with reasoning
MIT. See LICENSE.
If you've got a rule pattern I haven't caught, open a PR. If you've built an adapter for a CRM I don't have, open a PR for that too. If you think the whole approach is wrong, I'd love to hear why — open an issue.
- samwarren.io/projects/marketing-ops-agent — the production system this was extracted from (architecture only — the real version lives inside an employer's private repo)
- samwarren.io/projects/cascade — the broader operations layer this plugs into