Repository navigation
Conversation
|
Hi, thanks for the PR! This looks really cool, I think we definitely want this. Just a couple notes:
|
ea6db5f to
5c4c47e
Compare
|
Thanks @liquidsec — glad it's a fit! Retargeted to The Validated on the Totally fine if this lands in |
|
Had a chance to do a cursory review, there are a few things: The redactions:BBOT doesn't have any other cases where credentials harvested from a target are redacted in output. There are some instances where we redact our own api key secrets etc. But the philosophy here is, this is an offensive tool, many users will want to go on to actually use the harvested secrets. They'd also still be in output.json, so the potential usecase is definitely narrow at best and it adds complexity. The URL is also in the output, and re-fetch is trivial anyway. Missing Severity And Confidence in yara metadata:The custom severity and confidence fields get passed on to the eventual finding. In this case, we hope the regexes are good enough we can call them "confirmed". The severity we could debate about. However, in the absence of a severity it will be emitted as informational which is probably not what we want. RegexI suspect the provider endpoint stuff (looking for "api.openai.com") but not be specific enough, and i believe would just fire on any mentions of them Im not totally sure that the api key patterns wouldn't occasionally false positive but im not sure Process() overrideunderstand why you did this, but i think rather than do that override i'd almost rather split it into two separate excavate signatures, and then do some modifications to the base to support technology events. I'd need to think a little deeper into it first, but I can definitely help with that part. overallThe only other thing that makes me pause a little bit with this is the potential overlap with the trufflehog module. I'd want to do some more testing to see if trufflehog would pick this up... There's a larger question here which is the degree with which we want to outsource secret detection to trufflehog vs do it ourselves. I think what im leaning towards right now, pending actually getting my hands dirty w/this and testing, is leaving the API key detection to trufflehog and handle the detection of actual LLM endpoints with an excavate submodule. |
There was a problem hiding this comment.
@DevamShah, YARA-first is the right approach. @liquidsec already gave direction; I pulled the branch and confirmed his two "I'm not sure" items are real. Both reproduce end to end. Agreed on the rest.
Your two tests pass (5.44s), ruff clean.
|
I have read the CLA Document and I hereby sign the CLA You can retrigger this bot by commenting recheck in this Pull Request. Posted by the CLA Assistant Lite bot. |
|
@bb_top.md |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## dev #3235 +/- ##
=======================================
+ Coverage 90% 90% +1%
=======================================
Files 454 459 +5
Lines 47081 48561 +1480
=======================================
+ Hits 42323 43696 +1373
- Misses 4758 4865 +107 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
singlerider
left a comment
There was a problem hiding this comment.
@DevamShah Re-reviewed at 67e6cfa. Verified through the module harness, not isolated rule matching.
Prior four blockers resolved:
- Key detection removed from the diff.
$llamaindexand$langchainrequire import syntax.<a href="/static/llama_index.png">our llama, indexed</a>emits no TECHNOLOGY.- Provider endpoints anchored scheme plus host plus path. Bare hostnames in prose do not fire.
- Tests rewritten to exact-set positive and colliding-input negative.
Severity/confidence meta is inert as you stated: report() at excavate.py:487-495 injects those only when event_type == "FINDING".
🔴 Blocking, SSE rule fires on any event stream carrying delta or choices
Detail inline at excavate.py:1003. Generic progress-bar stream emits llm sse chat-completion stream. Negative fixture uses data: {"percent": 42}, which carries neither key, so it passes while the rule is broken.
Endpoint rule, not blocking
A page linking to a provider emits TECHNOLOGY:
ENDPOINT_PROBE=['llm chat-completion endpoint', 'openai api']
<p>See the <a href="https://api.openai.com/v1/chat/completions">API reference</a>.</p>
Web_Service_WSDL behaves identically on the same input shape:
WSDL_BASELINE=['HTTP response (body) contains a web service WSDL URL [https://vendor.example.net/api/foo.wsdl]']
Same property as the anchoring style you were pointed at. Not a blocker.
Difference for a separate issue, not this PR: WSDL emits FINDING with the matched URL in the description. TECHNOLOGY carries only a label, and asset_inventory.py:332 folds it into a per-host technology set.
@liquidsec, the process() override question from your first pass is still open.
…ked AI provider keys Adds a passive excavate submodule (AIApplicationExtractor) that fingerprints AI/LLM application surface (provider API endpoints, embedded client SDKs for LangChain / LlamaIndex / OpenAI) and flags leaked AI provider API keys (OpenAI, Anthropic, Google Gemini) from HTTP_RESPONSE bodies and JavaScript, emitting TECHNOLOGY and FINDING events. Reuses excavate's existing YARA pipeline; no new dependency. Detected secrets are redacted before emission. Includes two ModuleTestBase tests (positive + negative) mirroring the existing extractor test conventions. Signed-off-by: Devam Shah <devamshah91@gmail.com>
Addresses review feedback from @liquidsec and @singlerider. Remove the ai_provider_apikey rule and the redaction helper entirely. bbot/modules/trufflehog.py already watches HTTP_RESPONSE and RAW_TEXT -- the same inputs excavate sees -- so provider-key regexes here duplicated trufflehog with worse precision and no verification. The AIza pattern in particular is the generic Google API key format (Maps, Firebase, YouTube), published in page source by design and indistinguishable from a Gemini key. Anchor the remaining rules on structure rather than substrings: - provider endpoints now require a full URL (scheme + host + path), so prose naming a provider no longer fingerprints the target - /v1/chat/completions must appear inside a quoted string literal - LangChain and LlamaIndex now require import/require/instantiation syntax, mirroring the existing $openai_sdk shape Add severity and confidence meta to every rule. Rewrite both tests. The positive test compares the exact set of emitted TECHNOLOGY labels instead of grepping a substring. The negative test uses colliding inputs -- a Google Maps embed with an AIza key, a /static/llama_index.png href, prose naming both providers and the chat-completions path, and a non-LLM text/event-stream -- and asserts zero TECHNOLOGY events. Against the previous rules that test fails with four false positives. Signed-off-by: devamshah <devamshah91@gmail.com>
Addresses @singlerider's blocking review on ai_chat_streaming. $sse_chat_delta matched a bare top-level "delta" or "choices" key, and both are ordinary field names in unrelated event streams -- a price ticker publishing a delta move, or a poll publishing a flat choices list, each fingerprinted the host as running an LLM chat stream. Replace it with three markers that pin the structure a chat-completion frame actually has: - OpenAI nests delta inside the choices array: "choices":[{..."delta": - OpenAI labels the frame "object":"chat.completion.chunk" - Anthropic labels it "type":"content_block_delta" / "message_delta" The condition is now $sse_event_stream and any of ($sse_chat_*), so the content-type alone still cannot fire the rule. The negative fixture passed because it carried neither key. It now carries both, at the top level of non-LLM frames, so this regression fails the test instead of slipping through. Signed-off-by: devamshah <devamshah91@gmail.com>
67e6cfa to
61fbbce
Compare
singlerider
left a comment
There was a problem hiding this comment.
Re-reviewed at 61fbbce. Ran the module tests this time rather than matching rules in isolation, since you could not build the wheel locally.
TestExcavateAIApplicationPositive and TestExcavateAIApplicationNegative both pass on my machine. The SSE blocker is fixed. My exact ticker case no longer matches, nor does a flat choices list, while an OpenAI chunk, a choices-nested frame and an Anthropic content_block_delta all still do. Putting both keys at the top level of non-LLM frames in the negative fixture is what makes that stick, so the next person to loosen the rule will hear about it from CI.
Clearing my changes-requested.
Still open and not mine to settle: @liquidsec's process() override question from 2026-07-10. The override is still in the diff at excavate.py:1180. He floated splitting it into two signatures with base support for TECHNOLOGY events and offered to help shape it. That call is his.
singlerider
left a comment
There was a problem hiding this comment.
I looked one more time. I'd like for these excessive comments to please be stripped out. I think most that you added are not necessary.
Requested in review. The YARA rules and the tests already state what each signal matches and why the negative fixtures collide with it, so the comments restated the code. No rule, logic, or fixture changes.
singlerider
left a comment
There was a problem hiding this comment.
🟢 Re-reviewed at 97c638f. SSE rule now requires LLM-shaped frames, negative fixtures cover bare delta and choices. Both AIApplication tests pass.
|
@DevamShah approved, but the CLA check is still failing on your commits. Please sign by commenting exactly: I have read the CLA Document and I hereby sign the CLA |
|
After taking a closer look at this signatures, i think we need to pause and rethink this. I think if we are going to add signatures to detect AI integrations, we need to take the time to get some research to back up the efficacy of the signatures, and ensure we're not overlooking others as well. Many of the signatures are targetted towards source code, not things we'd expect to see deployed on a web application. I ran them against live traffic to check: across 413 random sites they produced zero false positives, which is good, but across 33 reachable sites that are AI vendors or ship LLM chat they produced exactly one detection, and that was on the chat-completions path. The import/require patterns get erased by webpack/vite/esbuild before anything ships, and the provider hostnames belong in backend code where the key lives, so we shouldn't expect to see them in the browser in the first place. The positive test only passes because the fixture is hand-written unbundled source instead of a real response. Before we settle on signatures again I'd want to look at an actual corpus of deployed AI-enabled apps and base them on what's observable in the response, rather than what the source would have looked like. |
|
Path forward: Lets back up each signature with an individual test, with a snippet procured from a real application using the technology found in a response, with a brief explanation of where it came from. This serves two purposes: One, it ensure each signature is well-documented so we can properly evaluate how to alter or remove it in the future as these technologies change. Two, if someone gets a detection with it, they will have some resource for understanding exactly what they've dicovered. |
Summary
Adds a passive
excavatesubmodule (AIApplicationExtractor) that fingerprints AI/LLM application surface — provider API endpoints, embedded client SDKs (LangChain / LlamaIndex / OpenAI) — and flags leaked AI provider API keys (OpenAI, Anthropic, Google Gemini) directly fromHTTP_RESPONSEbodies and JavaScript, emittingTECHNOLOGYandFINDINGevents.Problem / motivation
BBOT already excels at passively surfacing attack surface from HTTP responses, but it has no native recognition of the AI/LLM layer that now ships in most modern web front-ends. Two concrete gaps:
api.openai.com/v1/chat/completions,api.anthropic.com,generativelanguage.googleapis.com, Azure OpenAI deployments, and SSE chat streams, plus@langchain/*/llama_index/openaiSDK markers. These are high-value pivots for an assessor (prompt-injection sinks, server-side proxy endpoints, model/billing abuse) yet are invisible to current modules.sk-…,sk-proj-…), Anthropic (sk-ant-api03-…), and Google (AIza…) keys are regularly committed into front-end JS and inline<script>blobs. A leaked inference key is a direct financial-loss and data-exfiltration primitive.This reuses excavate's existing YARA pipeline (excavate already ships ~140 YARA references), so there is no new dependency or architectural change.
Change
AIApplicationExtractor(ExcavateRule)inbbot/modules/internal/excavate.py, placed alongside the other detection extractors (FunctionalityExtractor,SerializationExtractor,ErrorExtractor).ai_provider_endpoint— OpenAI / Anthropic / Google Generative AI / Azure OpenAI hosts +/v1/chat/completions.ai_chat_streaming— SSE chat stream, gated on bothtext/event-streamand an LLM-shapeddata: {... "delta"/"choices" ...}frame to avoid matching generic event streams.ai_client_library— LangChain (@langchain/…), LlamaIndex (llama[_-]?index), and OpenAI SDK import/instantiation markers.ai_provider_apikey— anchored, length-bounded key formats for OpenAI, OpenAI project, Anthropic, and Google keys.process()maps each matched YARA string identifier to either aTECHNOLOGYevent (endpoint/SDK fingerprint, requires a host) or aFINDINGevent (leaked key). Secrets are redacted (prefix…suffix) before they enter scan output rather than echoed in full.bbot/test/test_step_2/module_tests/test_module_excavate.pyfollowing the existingTestExcavateCSP/TestExcavateSerialization*conventions: a positive case (all techs + all three providers' keys) and a negative case (benign AI marketing copy + a Stripesk_live_key + "blockchain language" decoy) asserting zero false positives and that full secrets are never emitted.produced_eventson the parent module is intentionally left unchanged, consistent with the existing extractors (SerializationExtractor/CSPExtractoralready emitFINDING/DNS_NAMEwithout enumerating them there).Security rationale
sk-/sk-ant-/AIzakeys in client-delivered JS are a textbook instance; detecting them passively closes a common, high-impact reconnaissance gap.\b-anchored and length-bounded, provider key namespaces are mutually exclusive (hyphenatedsk-ant-/sk-proj-cannot match the 48-char legacy OpenAI pattern), the SSE rule requires two independent signals, and detected secrets are redacted before emission.Testing / validation
Validated against a fresh clone with
yara-python4.5.2 (the version BBOT pins) and the fullbbot/testharness on Python 3.12:bbot/testpytest run:TECHNOLOGYevents (OpenAI/Anthropic/Gemini API, chat-completion endpoint, LangChain, OpenAI SDK) andFINDINGevents for the three leaked provider keys, and that full secrets are never echoed. The negative case asserts zero AIFINDING/TECHNOLOGYon benign content (Stripesk_live_, AI marketing prose, "blockchain language model").ruff check→ "All checks passed!";ruff format --check→ clean (line-length 119).The submodule follows the same shape as the existing
SerializationExtractor/FunctionalityExtractorextractors, so it carries no standalone doc file (module docs are generated from metadata).AI Use Disclosure (per
docs/contribution.md)Claude Opus, used extensively — closer to the autonomous end than a back-and-forth, with me setting the direction and reviewing the result. That applies to the original submission and to the 67e6cfa revision.
Flagging one thing directly, since your policy says not to let the AI edit tests: the tests in 67e6cfa were rewritten, because that was the specific change @singlerider asked for. I've checked them by hand — the negative test is built from his colliding inputs and I confirmed it fails with four FPs against the old rules before trusting it, and the positive test now compares the full nine-label set instead of a substring grep. But it's your rule and your call, so if you'd rather review those two functions separately or have me redo them differently, say so.