Skip to content

excavate: add AIApplicationExtractor for LLM endpoints, SDKs, and leaked AI provider keys - #3235

Open
DevamShah wants to merge 4 commits into
blacklanternsecurity:devfrom
DevamShah:excavate-ai-application-extractor
Open

DevamShah wants to merge 4 commits into
blacklanternsecurity:devfrom
DevamShah:excavate-ai-application-extractor

Conversation

@DevamShah

@DevamShah DevamShah commented Jun 23, 2026 •

Copy link
Copy Markdown

Summary

Adds a passive excavate submodule (AIApplicationExtractor) that fingerprints AI/LLM application surface — provider API endpoints, embedded client SDKs (LangChain / LlamaIndex / OpenAI) — and flags leaked AI provider API keys (OpenAI, Anthropic, Google Gemini) directly from HTTP_RESPONSE bodies and JavaScript, emitting TECHNOLOGY and FINDING events.

Problem / motivation

BBOT already excels at passively surfacing attack surface from HTTP responses, but it has no native recognition of the AI/LLM layer that now ships in most modern web front-ends. Two concrete gaps:

  1. No AI attack-surface fingerprinting. Client-side bundles routinely hard-code calls to api.openai.com/v1/chat/completions, api.anthropic.com, generativelanguage.googleapis.com, Azure OpenAI deployments, and SSE chat streams, plus @langchain/* / llama_index / openai SDK markers. These are high-value pivots for an assessor (prompt-injection sinks, server-side proxy endpoints, model/billing abuse) yet are invisible to current modules.
  2. No detection of leaked AI provider credentials. OpenAI (sk-…, sk-proj-…), Anthropic (sk-ant-api03-…), and Google (AIza…) keys are regularly committed into front-end JS and inline <script> blobs. A leaked inference key is a direct financial-loss and data-exfiltration primitive.

This reuses excavate's existing YARA pipeline (excavate already ships ~140 YARA references), so there is no new dependency or architectural change.

Change

  • New AIApplicationExtractor(ExcavateRule) in bbot/modules/internal/excavate.py, placed alongside the other detection extractors (FunctionalityExtractor, SerializationExtractor, ErrorExtractor).
  • Four YARA rules, compiled into excavate's existing combined ruleset:
    • ai_provider_endpoint — OpenAI / Anthropic / Google Generative AI / Azure OpenAI hosts + /v1/chat/completions.
    • ai_chat_streaming — SSE chat stream, gated on both text/event-stream and an LLM-shaped data: {... "delta"/"choices" ...} frame to avoid matching generic event streams.
    • ai_client_library — LangChain (@langchain/…), LlamaIndex (llama[_-]?index), and OpenAI SDK import/instantiation markers.
    • ai_provider_apikey — anchored, length-bounded key formats for OpenAI, OpenAI project, Anthropic, and Google keys.
  • process() maps each matched YARA string identifier to either a TECHNOLOGY event (endpoint/SDK fingerprint, requires a host) or a FINDING event (leaked key). Secrets are redacted (prefix…suffix) before they enter scan output rather than echoed in full.
  • Two new test classes in bbot/test/test_step_2/module_tests/test_module_excavate.py following the existing TestExcavateCSP / TestExcavateSerialization* conventions: a positive case (all techs + all three providers' keys) and a negative case (benign AI marketing copy + a Stripe sk_live_ key + "blockchain language" decoy) asserting zero false positives and that full secrets are never emitted.

produced_events on the parent module is intentionally left unchanged, consistent with the existing extractors (SerializationExtractor/CSPExtractor already emit FINDING/DNS_NAME without enumerating them there).

Security rationale

  • OWASP LLM Top 10 (2025): maps to LLM02 — Sensitive Information Disclosure (leaked provider keys) and LLM10 — Unbounded Consumption (exposed inference endpoints enabling model/billing abuse). Surfacing the endpoints also locates the LLM01 — Prompt Injection sinks for follow-on testing.
  • CWE-798 (Use of Hard-coded Credentials) and CWE-312 (Cleartext Storage of Sensitive Information): leaked sk-/sk-ant-/AIza keys in client-delivered JS are a textbook instance; detecting them passively closes a common, high-impact reconnaissance gap.
  • MITRE ATT&CK T1552.001 (Unsecured Credentials: Credentials in Files): client bundles are exactly the kind of artifact where these credentials end up.
  • False-positive discipline is treated as a first-class requirement: every key pattern is \b-anchored and length-bounded, provider key namespaces are mutually exclusive (hyphenated sk-ant-/sk-proj- cannot match the 48-char legacy OpenAI pattern), the SSE rule requires two independent signals, and detected secrets are redacted before emission.

Testing / validation

Validated against a fresh clone with yara-python 4.5.2 (the version BBOT pins) and the full bbot/test harness on Python 3.12:

  1. In-harness tests pass. Both new classes pass under the standard bbot/test pytest run:
    pytest bbot/test/test_step_2/module_tests/test_module_excavate.py::TestExcavateAIApplicationPositive \
           bbot/test/test_step_2/module_tests/test_module_excavate.py::TestExcavateAIApplicationNegative
    # 2 passed
    
    The positive case asserts the expected TECHNOLOGY events (OpenAI/Anthropic/Gemini API, chat-completion endpoint, LangChain, OpenAI SDK) and FINDING events for the three leaked provider keys, and that full secrets are never echoed. The negative case asserts zero AI FINDING/TECHNOLOGY on benign content (Stripe sk_live_, AI marketing prose, "blockchain language model").
  2. Lint/format: ruff check → "All checks passed!"; ruff format --check → clean (line-length 119).

The submodule follows the same shape as the existing SerializationExtractor/FunctionalityExtractor extractors, so it carries no standalone doc file (module docs are generated from metadata).


AI Use Disclosure (per docs/contribution.md)

Claude Opus, used extensively — closer to the autonomous end than a back-and-forth, with me setting the direction and reviewing the result. That applies to the original submission and to the 67e6cfa revision.

Flagging one thing directly, since your policy says not to let the AI edit tests: the tests in 67e6cfa were rewritten, because that was the specific change @singlerider asked for. I've checked them by hand — the negative test is built from his colliding inputs and I confirmed it fails with four FPs against the old rules before trusting it, and the positive test now compares the full nine-label set instead of a substring grep. But it's your rule and your call, so if you'd rather review those two functions separately or have me redo them differently, say so.

@liquidsec

Copy link
Copy Markdown
Contributor

Hi, thanks for the PR! This looks really cool, I think we definitely want this. Just a couple notes:

  • We are weeks away from 3.0 release, currently on dev. Can you retarget to dev? There is a VERY significant difference between 2.8 and 3.0, i'd anticipate conflicts.
  • Also as a result of being so close to release, I am not sure if it will make it in 3.0 since we are essentially freezing new capabilities until after that. Most likely it would go in dev after 3.0 release, and ultimately be released in 3.1.

@liquidsec
liquidsec changed the base branch from stable to dev June 23, 2026 21:15
@DevamShah
DevamShah force-pushed the excavate-ai-application-extractor branch from ea6db5f to 5c4c47e Compare June 24, 2026 04:33
@DevamShah

Copy link
Copy Markdown
Author

Thanks @liquidsec — glad it's a fit! Retargeted to dev and rebased the branch onto current dev (3.0).

The ExcavateRule contract was unchanged between 2.8 and 3.0, so the module itself applied cleanly with no logic changes. One 3.0 behavior change did surface in testing: the TECHNOLOGY event now lowercases its technology field in _sanitize_data, so I adapted the positive/negative test assertions to compare case-insensitively (the module still emits the same readable labels — core just normalizes the case). No assertions were weakened.

Validated on the dev base: the full test_module_excavate.py suite passes (52 passed, including the two new AIApplication positive/negative tests), and ruff check/ruff format are clean.

Totally fine if this lands in dev post-3.0 / ships in 3.1 — no rush on my end. Happy to make any further changes you'd like.

@liquidsec liquidsec added this to the BBOT 3.1 - violent_barbara milestone Jul 10, 2026
@liquidsec liquidsec self-assigned this Jul 10, 2026
@liquidsec

liquidsec commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Had a chance to do a cursory review, there are a few things:

The redactions:

BBOT doesn't have any other cases where credentials harvested from a target are redacted in output. There are some instances where we redact our own api key secrets etc. But the philosophy here is, this is an offensive tool, many users will want to go on to actually use the harvested secrets. They'd also still be in output.json, so the potential usecase is definitely narrow at best and it adds complexity. The URL is also in the output, and re-fetch is trivial anyway.

Missing Severity And Confidence in yara metadata:

The custom severity and confidence fields get passed on to the eventual finding. In this case, we hope the regexes are good enough we can call them "confirmed". The severity we could debate about. However, in the absence of a severity it will be emitted as informational which is probably not what we want.

Regex

I suspect the provider endpoint stuff (looking for "api.openai.com") but not be specific enough, and i believe would just fire on any mentions of them

Im not totally sure that the api key patterns wouldn't occasionally false positive but im not sure

Process() override

understand why you did this, but i think rather than do that override i'd almost rather split it into two separate excavate signatures, and then do some modifications to the base to support technology events. I'd need to think a little deeper into it first, but I can definitely help with that part.

overall

The only other thing that makes me pause a little bit with this is the potential overlap with the trufflehog module. I'd want to do some more testing to see if trufflehog would pick this up...

There's a larger question here which is the degree with which we want to outsource secret detection to trufflehog vs do it ourselves.

I think what im leaning towards right now, pending actually getting my hands dirty w/this and testing, is leaving the API key detection to trufflehog and handle the detection of actual LLM endpoints with an excavate submodule.

@singlerider singlerider left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@DevamShah, YARA-first is the right approach. @liquidsec already gave direction; I pulled the branch and confirmed his two "I'm not sure" items are real. Both reproduce end to end. Agreed on the rest.

Your two tests pass (5.44s), ruff clean.

Comment thread bbot/modules/internal/excavate.py Outdated
Comment thread bbot/modules/internal/excavate.py Outdated
Comment thread bbot/modules/internal/excavate.py Outdated
Comment thread bbot/test/test_step_2/module_tests/test_module_excavate.py Outdated
@github-actions

Copy link
Copy Markdown
Contributor


Thank you for your submission, we really appreciate it. Like many open-source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution. You can sign the CLA by just posting a Pull Request Comment same as the below format.


I have read the CLA Document and I hereby sign the CLA


You can retrigger this bot by commenting recheck in this Pull Request. Posted by the CLA Assistant Lite bot.

@DevamShah

DevamShah commented Aug 28, 2026 •

Copy link
Copy Markdown
Author

@bb_top.md

@codecov

codecov Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.36842% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 90%. Comparing base (0c510b1) to head (97c638f).
⚠️ Report is 27 commits behind head on dev.

Files with missing lines Patch % Lines
bbot/modules/internal/excavate.py 95% 1 Missing ⚠️
Additional details and impacted files
@@           Coverage Diff           @@
##             dev   #3235     +/-   ##
=======================================
+ Coverage     90%     90%     +1%     
=======================================
  Files        454     459      +5     
  Lines      47081   48561   +1480     
=======================================
+ Hits       42323   43696   +1373     
- Misses      4758    4865    +107     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@singlerider singlerider left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@DevamShah Re-reviewed at 67e6cfa. Verified through the module harness, not isolated rule matching.

Prior four blockers resolved:

  • Key detection removed from the diff.
  • $llamaindex and $langchain require import syntax. <a href="/static/llama_index.png">our llama, indexed</a> emits no TECHNOLOGY.
  • Provider endpoints anchored scheme plus host plus path. Bare hostnames in prose do not fire.
  • Tests rewritten to exact-set positive and colliding-input negative.

Severity/confidence meta is inert as you stated: report() at excavate.py:487-495 injects those only when event_type == "FINDING".

🔴 Blocking, SSE rule fires on any event stream carrying delta or choices

Detail inline at excavate.py:1003. Generic progress-bar stream emits llm sse chat-completion stream. Negative fixture uses data: {"percent": 42}, which carries neither key, so it passes while the rule is broken.

Endpoint rule, not blocking

A page linking to a provider emits TECHNOLOGY:

ENDPOINT_PROBE=['llm chat-completion endpoint', 'openai api']
  <p>See the <a href="https://api.openai.com/v1/chat/completions">API reference</a>.</p>

Web_Service_WSDL behaves identically on the same input shape:

WSDL_BASELINE=['HTTP response (body) contains a web service WSDL URL [https://vendor.example.net/api/foo.wsdl]']

Same property as the anchoring style you were pointed at. Not a blocker.

Difference for a separate issue, not this PR: WSDL emits FINDING with the matched URL in the description. TECHNOLOGY carries only a label, and asset_inventory.py:332 folds it into a per-host technology set.

@liquidsec, the process() override question from your first pass is still open.

Comment thread bbot/modules/internal/excavate.py Outdated
…ked AI provider keys

Adds a passive excavate submodule (AIApplicationExtractor) that fingerprints
AI/LLM application surface (provider API endpoints, embedded client SDKs for
LangChain / LlamaIndex / OpenAI) and flags leaked AI provider API keys
(OpenAI, Anthropic, Google Gemini) from HTTP_RESPONSE bodies and JavaScript,
emitting TECHNOLOGY and FINDING events. Reuses excavate's existing YARA
pipeline; no new dependency. Detected secrets are redacted before emission.

Includes two ModuleTestBase tests (positive + negative) mirroring the existing
extractor test conventions.

Signed-off-by: Devam Shah <devamshah91@gmail.com>
Addresses review feedback from @liquidsec and @singlerider.

Remove the ai_provider_apikey rule and the redaction helper entirely.
bbot/modules/trufflehog.py already watches HTTP_RESPONSE and RAW_TEXT --
the same inputs excavate sees -- so provider-key regexes here duplicated
trufflehog with worse precision and no verification. The AIza pattern in
particular is the generic Google API key format (Maps, Firebase, YouTube),
published in page source by design and indistinguishable from a Gemini key.

Anchor the remaining rules on structure rather than substrings:
- provider endpoints now require a full URL (scheme + host + path), so
  prose naming a provider no longer fingerprints the target
- /v1/chat/completions must appear inside a quoted string literal
- LangChain and LlamaIndex now require import/require/instantiation
  syntax, mirroring the existing $openai_sdk shape

Add severity and confidence meta to every rule.

Rewrite both tests. The positive test compares the exact set of emitted
TECHNOLOGY labels instead of grepping a substring. The negative test uses
colliding inputs -- a Google Maps embed with an AIza key, a
/static/llama_index.png href, prose naming both providers and the
chat-completions path, and a non-LLM text/event-stream -- and asserts
zero TECHNOLOGY events. Against the previous rules that test fails with
four false positives.

Signed-off-by: devamshah <devamshah91@gmail.com>
Addresses @singlerider's blocking review on ai_chat_streaming.

$sse_chat_delta matched a bare top-level "delta" or "choices" key, and
both are ordinary field names in unrelated event streams -- a price ticker
publishing a delta move, or a poll publishing a flat choices list, each
fingerprinted the host as running an LLM chat stream.

Replace it with three markers that pin the structure a chat-completion
frame actually has:
- OpenAI nests delta inside the choices array: "choices":[{..."delta":
- OpenAI labels the frame "object":"chat.completion.chunk"
- Anthropic labels it "type":"content_block_delta" / "message_delta"

The condition is now $sse_event_stream and any of ($sse_chat_*), so the
content-type alone still cannot fire the rule.

The negative fixture passed because it carried neither key. It now carries
both, at the top level of non-LLM frames, so this regression fails the test
instead of slipping through.

Signed-off-by: devamshah <devamshah91@gmail.com>
@DevamShah
DevamShah force-pushed the excavate-ai-application-extractor branch from 67e6cfa to 61fbbce Compare September 18, 2026 06:10

@singlerider singlerider left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed at 61fbbce. Ran the module tests this time rather than matching rules in isolation, since you could not build the wheel locally.

TestExcavateAIApplicationPositive and TestExcavateAIApplicationNegative both pass on my machine. The SSE blocker is fixed. My exact ticker case no longer matches, nor does a flat choices list, while an OpenAI chunk, a choices-nested frame and an Anthropic content_block_delta all still do. Putting both keys at the top level of non-LLM frames in the negative fixture is what makes that stick, so the next person to loosen the rule will hear about it from CI.

Clearing my changes-requested.

Still open and not mine to settle: @liquidsec's process() override question from 2026-07-10. The override is still in the diff at excavate.py:1180. He floated splitting it into two signatures with base support for TECHNOLOGY events and offered to help shape it. That call is his.

@singlerider singlerider left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I looked one more time. I'd like for these excessive comments to please be stripped out. I think most that you added are not necessary.

@liquidsec liquidsec removed this from the BBOT 3.1 - violent_barbara milestone Oct 1, 2026
Requested in review. The YARA rules and the tests already state what
each signal matches and why the negative fixtures collide with it, so
the comments restated the code. No rule, logic, or fixture changes.

@singlerider singlerider left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Re-reviewed at 97c638f. SSE rule now requires LLM-shaped frames, negative fixtures cover bare delta and choices. Both AIApplication tests pass.

@singlerider

Copy link
Copy Markdown
Contributor

@DevamShah approved, but the CLA check is still failing on your commits. Please sign by commenting exactly:

I have read the CLA Document and I hereby sign the CLA

@liquidsec

liquidsec commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

After taking a closer look at this signatures, i think we need to pause and rethink this. I think if we are going to add signatures to detect AI integrations, we need to take the time to get some research to back up the efficacy of the signatures, and ensure we're not overlooking others as well.

Many of the signatures are targetted towards source code, not things we'd expect to see deployed on a web application. I ran them against live traffic to check: across 413 random sites they produced zero false positives, which is good, but across 33 reachable sites that are AI vendors or ship LLM chat they produced exactly one detection, and that was on the chat-completions path.

The import/require patterns get erased by webpack/vite/esbuild before anything ships, and the provider hostnames belong in backend code where the key lives, so we shouldn't expect to see them in the browser in the first place. The positive test only passes because the fixture is hand-written unbundled source instead of a real response. Before we settle on signatures again I'd want to look at an actual corpus of deployed AI-enabled apps and base them on what's observable in the response, rather than what the source would have looked like.

@liquidsec liquidsec left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Signatures need more research. Many of them will never fire because they are targeted towards source code, not web content.

@liquidsec

Copy link
Copy Markdown
Contributor

Path forward: Lets back up each signature with an individual test, with a snippet procured from a real application using the technology found in a response, with a brief explanation of where it came from.

This serves two purposes: One, it ensure each signature is well-documented so we can properly evaluate how to alter or remove it in the future as these technologies change. Two, if someone gets a detection with it, they will have some resource for understanding exactly what they've dicovered.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants