Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@
- Added direct DSH profile plugin discovery and DSH-specific risk scanning to the standard `agentguard checkup` workflow.

### Fixed
- Preserved bounded, validated protected-file summaries such as `cat .env` and `cat .ssh/id_ed25519.pub` in native-hook audit and Cloud event records while continuing to suppress absolute paths, extra arguments, command tails, and unsafe filenames.
- Kept routine model-request records and all model-response records out of Cloud event ingest while preserving redacted local audit; confirmed PII-bearing requests are reported only when they were not stopped before model egress.
- Fixed Windows and system-cron patrols to run the existing eight-check `agentguard checkup --json` flow while keeping SessionStart lightweight.
- Improved DSH subscription cleanup and artifact discovery, and made system cron status failures explicit.
- Fixed `checkup` to recursively scan plugins referenced by DSH bundles, wait for all DSH scans before report generation, and include per-plugin results in JSON and HTML reports.
Expand Down
2 changes: 1 addition & 1 deletion docs/claude-code.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,6 @@ An `@file` path without an explicit exact deny is `unsupported` and is not count

## Privacy and coverage

Raw prompts, command arguments, tool output, failure text, configuration content, credentials, and PII are evaluated locally. Native Claude hook audit and Cloud records replace hook content with `[LOCAL_ONLY_LLM_CONTENT]` and retain only bounded rule IDs, masks, counts, action IDs, coverage, and enforcement metadata. Hook stdout/stderr never echoes raw input.
Raw prompts, command arguments, tool output, failure text, configuration content, credentials, and PII are evaluated locally. Native Claude hook audit and Cloud records replace raw hook content with `[LOCAL_ONLY_LLM_CONTENT]` and retain only bounded rule IDs, masks, counts, action IDs, coverage, and enforcement metadata. For protected-file tool actions, they may retain only a bounded, validated explanation such as `cat .env` or `cat .ssh/id_ed25519.pub`, without the absolute path, remaining arguments, or command tail. Hook stdout/stderr never echoes raw input.

Claude Code provides strong staged prompt/tool/context protection, but it does not expose the final model HTTP destination, Authorization credential, complete assembled payload, full response, all retries/fallbacks, or auxiliary model calls. Model-transport rules therefore remain `partial` or `unsupported`; AgentGuard does not claim complete model-traffic interception.
9 changes: 6 additions & 3 deletions docs/codex.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,9 +77,12 @@ Malformed hook JSON and catchable evaluator/process errors cause the installed
wrapper to exit `2` with a generic, non-sensitive reason. If Codex forcibly
terminates the hook or the process cannot return an exit code, Codex's host
behavior applies; this integration does not claim those timeouts fail closed.
Hook stdout, stderr, audit, and Cloud events contain only bounded rule ids,
risk, action ids, coverage facts, and redacted reasons—not prompt text, tool
output, credentials, or PII.
Hook stdout and stderr contain only bounded rule ids, risk, action ids, coverage
facts, and redacted reasons. Audit and Cloud events do not contain prompt text,
tool output, credentials, or PII; a protected-file tool action may retain only
a bounded, validated explanation such as `cat .env` or
`cat .ssh/id_ed25519.pub`, without its absolute path, remaining arguments, or
command tail.

## Coverage and known gaps

Expand Down
6 changes: 3 additions & 3 deletions docs/openclaw.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,9 +47,9 @@ AgentGuard registers only OpenClaw's existing public plugin hooks:
| Hook | AgentGuard behavior | Boundary |
| --- | --- | --- |
| `before_agent_run` | Blocking gate over the initial prompt, loaded history, and system prompt | Runs once at the supported run boundary; it is not a gate for every model call |
| `llm_input` | Redacted semantic request audit | Observer only |
| `llm_output` | Redacted partial response audit | Observer only |
| `model_call_started` / `model_call_ended` | Correlates `runId` / `callId` and records provider, model, and supplied byte statistics | Observer only |
| `llm_input` | Redacted local semantic request audit; only confirmed, non-blocked PII egress is Cloud-reportable | Observer only |
| `llm_output` | Redacted local partial response audit; never Cloud-reportable | Observer only |
| `model_call_started` / `model_call_ended` | Locally correlates `runId` / `callId` and records provider, model, and supplied byte statistics; routine records are never Cloud-reportable | Observer only |
| `before_tool_call` | Blocks dangerous commands and sensitive actions or returns native OpenClaw approval | Blocking |
| `after_tool_call` | Audits tool outcomes and visible response-poisoning evidence | Observer only |

Expand Down
22 changes: 16 additions & 6 deletions docs/privacy-boundary.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,10 +11,10 @@ AgentGuard OSS protects your machine without requiring a Cloud account.
- Local audit file at `~/.agentguard/audit.jsonl`
- Cached policy at `~/.agentguard/policy-cache.json`

LLM request and response content is evaluated only in local process memory. The
persisted and Cloud-facing form replaces that content with
`[LOCAL_ONLY_LLM_CONTENT]`; it retains only bounded lifecycle and coverage
facts.
LLM request and response content is evaluated only in local process memory.
Redacted request and response observations may be written to the local audit
file, but routine model-request records and all model-response records are not
uploaded to Cloud.

Cloud runtime policy responses use `schemaVersion: 1`; older responses without
that field are treated as version 1 and normalized locally. Cloud audit and
Expand All @@ -24,10 +24,14 @@ Unknown fields are ignored rather than treated as local enforcement facts.

## Sent to Cloud when connected

Only redacted runtime audit previews are uploaded by default:
Only redacted, Cloud-eligible runtime audit previews are uploaded by default:

- `sessionId`, `agentHost`, `actionType`, `toolName`
- Redacted `input` preview, capped at 2,000 characters
- For native tool hooks, raw arguments remain local. Protected-file access may
include only a bounded, validated target summary such as `cat .env` or
`cat .ssh/id_ed25519.pub`; absolute paths, additional arguments, and command
tails are omitted.
- Decision, risk score, risk level, reasons, and policy version
- Lifecycle stage, coverage level, enforcement status, missing-fact names, and
request correlation IDs
Expand All @@ -36,6 +40,8 @@ Only redacted runtime audit previews are uploaded by default:
- Credential kind and presence only (`api_key`, `oauth`, `aws`, `ambient`,
`none`, or `unknown`); never a credential value, Authorization header, or
reversible digest
- For model traffic, only a confirmed PII-bearing request that was not stopped
before egress; routine requests and every response stay local
- PII category/count summaries and masked evidence; never raw matches

PII summaries use only an allowlisted category name, per-category count, and a
Expand Down Expand Up @@ -67,7 +73,11 @@ Cloud endpoints also apply server-side redaction, but clients should not rely on

## Offline behavior

If Cloud is unreachable, AgentGuard continues local enforcement and spools redacted audit events for later retry. It must never fail open for local `block` decisions.
If Cloud is unreachable, AgentGuard continues local enforcement and spools
Cloud-eligible redacted audit events for later retry. Routine model requests and
model responses are filtered again at the Cloud client boundary, including
when an older spool is drained. It must never fail open for local `block`
decisions.

Local policies are normalized when loaded, so caches written before the LLM
privacy fields existed inherit the bundled privacy defaults. Cloud availability
Expand Down
5 changes: 4 additions & 1 deletion src/cloud/client.ts
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ import {
isClaudeNativeHookAction,
isCodexNativeHookAction,
nativeHookSafeMetadata,
shouldReportAuditEventToCloud,
} from '../runtime/audit.js';
import type { Advisory, SelfCheckMatch } from '../feed/types.js';

Expand Down Expand Up @@ -86,10 +87,12 @@ export class AgentGuardCloudClient {

async ingestEvents(events: RuntimeAuditEvent[]): Promise<void> {
this.requireCredential();
const reportableEvents = events.filter((event) => shouldReportAuditEventToCloud(event));
if (reportableEvents.length === 0) return;
await this.request('/api/v1/events/ingest', {
method: 'POST',
body: JSON.stringify({
events: events.map((event) => buildCloudAuditEvent(event)),
events: reportableEvents.map((event) => buildCloudAuditEvent(event)),
}),
});
}
Expand Down
144 changes: 142 additions & 2 deletions src/runtime/audit.ts
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,11 @@ export function buildAuditEvent(event: RuntimeAuditEvent): RuntimeAuditEvent {
agentHost: event.agentHost,
actionType: event.actionType,
toolName: redactPreview(event.toolName, 160),
input: isLlmTrafficEvent(event) || nativeHook ? '[LOCAL_ONLY_LLM_CONTENT]' : redactPreview(event.input),
input: isLlmTrafficEvent(event)
? '[LOCAL_ONLY_LLM_CONTENT]'
: nativeHook
? nativeHookAuditInput(event)
: redactPreview(event.input),
decision: event.decision,
policyDecision: event.policyDecision,
riskScore: clampRiskScore(event.riskScore),
Expand All @@ -36,6 +40,34 @@ export function buildAuditEvent(event: RuntimeAuditEvent): RuntimeAuditEvent {
};
}

/**
* Keep model traffic private while still reporting confirmed PII egress.
*
* Model responses never leave the machine. Routine model requests also stay
* local; a request is Cloud-reportable only when local evaluation found PII
* and the request was not stopped at a blocking gate.
*/
export function shouldReportAuditEventToCloud(event: RuntimeAuditEvent): boolean {
if (event.actionType === 'llm_response') return false;
if (event.actionType !== 'llm_request') return true;

const piiDetected = Boolean(
event.privacySummary
&& Number.isSafeInteger(event.privacySummary.valueCount)
&& event.privacySummary.valueCount > 0
&& event.privacySummary.categories.some((item) => (
PII_CATEGORIES.has(item.category)
&& Number.isSafeInteger(item.count)
&& item.count > 0
))
);
if (!piiDetected) return false;

const stoppedBeforeEgress = event.canBlockCurrentAction !== false
&& (event.decision === 'block' || event.decision === 'require_approval');
return !stoppedBeforeEgress;
}

export function isCodexNativeHookAction(action: Pick<RuntimeAction, 'agentHost' | 'metadata'>): boolean {
return action.agentHost === 'codex' && CODEX_HOOK_EVENTS.has(action.metadata?.codexHookEvent);
}
Expand Down Expand Up @@ -86,13 +118,114 @@ function codexSafeReasons(reasons: PolicyReason[]): PolicyReason[] {
return reasons.slice(0, 20).map((reason) => ({
code: SAFE_RULE_ID.test(reason.code) ? reason.code : 'POLICY',
severity: CODEX_SEVERITIES.has(reason.severity) ? reason.severity as RuntimeSeverity : 'info',
title: 'Policy rule matched',
title: reason.code === 'SECRET_ACCESS' ? 'Protected path access' : 'Policy rule matched',
description: '[REDACTED]',
evidence: reason.evidence === undefined ? undefined : '[REDACTED]',
remediation: reason.remediation === undefined ? undefined : '[REDACTED]',
}));
}

/**
* Preserve a minimal explanation for protected-file access
* without exporting the raw native-hook command, absolute path, or arguments.
*/
function nativeHookAuditInput(event: RuntimeAuditEvent): string {
if (!event.reasons.some((reason) => reason.code === 'SECRET_ACCESS')) {
return '[LOCAL_ONLY_LLM_CONTENT]';
}

const protectedFile = safeProtectedFileReference(event);
if (!protectedFile) return '[LOCAL_ONLY_LLM_CONTENT]';

if (event.actionType === 'shell') {
const command = safeShellCommandName(event.input);
return `${command ?? 'access'} ${protectedFile}`;
}
if (event.actionType === 'file_read') return `read ${protectedFile}`;
if (event.actionType === 'file_write') return `write ${protectedFile}`;
return `access ${protectedFile}`;
}

function safeProtectedFileReference(event: RuntimeAuditEvent): string | undefined {
const input = event.input;
const ssh = input.match(/\.ssh[\\/]([A-Za-z0-9._-]{1,128})(?=$|[\s"';&|)\]},:])/i)?.[1];
if (ssh && isSafeProtectedFileName(ssh)) return `.ssh/${ssh}`;

const aws = input.match(/\.aws[\\/]([A-Za-z0-9._-]{1,128})(?=$|[\s"';&|)\]},:])/i)?.[1];
if (aws && isSafeProtectedFileName(aws)) return `.aws/${aws}`;

const environment = input.match(
/(?:^|[\s"'=([{,:])(?:[A-Za-z]:[\\/])?[\\/]?(?:[A-Za-z0-9_~.-]+[\\/])*(\.env(?:\.[A-Za-z0-9_-]{1,64})?)(?=$|[\s"';&|)\]},:])/i
)?.[1];
if (environment && isSafeProtectedFileName(environment)) return environment;

const genericTarget = safeGenericProtectedTarget(event);
if (genericTarget) return genericTarget;

const credentials = input.match(
/(?:^|[\\/\s"'=([{,:])(credentials[A-Za-z0-9._-]{0,96})(?=$|[\s"';&|)\]},:])/i
)?.[1];
if (credentials && isSafeProtectedFileName(credentials)) return credentials;

const namedSecret = input.match(
/(?:^|[\\/\s"'=([{,:])([A-Za-z0-9._-]{0,96}(?:private-key|seed)[A-Za-z0-9._-]{0,96})(?=$|[\s"';&|)\]},:])/i
)?.[1];
if (namedSecret && isSafeProtectedFileName(namedSecret)) return namedSecret;

for (const reason of event.reasons) {
if (reason.code !== 'SECRET_ACCESS' || !reason.evidence || /[*?\[\]]/.test(reason.evidence)) continue;
const exactName = reason.evidence.split(/[\\/]/).at(-1);
if (event.input.includes(reason.evidence) && exactName && isSafeProtectedFileName(exactName)) return exactName;
}
return undefined;
}

function safeGenericProtectedTarget(event: RuntimeAuditEvent): string | undefined {
if (event.actionType === 'file_read') return safePathBasename(event.input);

if (event.actionType === 'file_write') {
const path = event.input.match(
/["'](?:file_path|filePath|path|target)["']\s*:\s*["']([^"']+)["']/
)?.[1];
return path ? safePathBasename(path) : undefined;
}

if (event.actionType !== 'shell') return undefined;
const command = safeShellCommandName(event.input);
if (!command || !SINGLE_TARGET_FILE_COMMANDS.has(command)) return undefined;
const firstCommand = event.input.split(/[;&|]/, 1)[0] ?? '';
const tokens = firstCommand.match(/"[^"]*"|'[^']*'|[^\s]+/g) ?? [];
for (let index = tokens.length - 1; index >= 0; index -= 1) {
const token = tokens[index]!.replace(/^["']|["']$/g, '');
if (token.startsWith('-') || /^\d+$/.test(token)) continue;
const name = safePathBasename(token);
if (name && name.toLowerCase() !== command) return name;
}
return undefined;
}

function safePathBasename(value: string): string | undefined {
const normalized = value.trim().replace(/^["']|["']$/g, '').replace(/[\\/]+$/, '');
const name = normalized.split(/[\\/]/).at(-1);
return name && isSafeProtectedFileName(name) ? name : undefined;
}

function isSafeProtectedFileName(value: string): boolean {
return value !== '.'
&& value !== '..'
&& value.length <= 128
&& /^[A-Za-z0-9._-]+$/.test(value)
&& redactPreview(value, 128) === value;
}

function safeShellCommandName(input: string): string | undefined {
const match = input.match(
/^\s*(?:(?:sudo|env)\s+)?(?:[A-Za-z_][A-Za-z0-9_]*=[^\s]+\s+)*(?:\/[^\s/]+\/)*([A-Za-z][A-Za-z0-9_-]*)\b/
);
const command = match?.[1]?.toLowerCase();
return command && SAFE_FILE_COMMANDS.has(command) ? command : undefined;
}

function codexSafePrivacyRule(value: unknown): RuntimePrivacyRuleEvaluation | null {
if (!value || typeof value !== 'object' || Array.isArray(value)) return null;
const rule = value as Record<string, unknown>;
Expand Down Expand Up @@ -150,6 +283,13 @@ const CODEX_DECISIONS = new Set<unknown>(['allow', 'warn', 'require_approval', '
const CODEX_SEVERITIES = new Set<unknown>(['info', 'low', 'medium', 'high', 'critical']);
const SAFE_RULE_ID = /^[A-Z][A-Z0-9_]{0,63}$/;
const SAFE_ID = /^[A-Za-z0-9_-]{1,160}$/;
const SAFE_FILE_COMMANDS = new Set([
'cat', 'head', 'tail', 'less', 'more', 'grep', 'sed', 'awk',
'cp', 'mv', 'rm', 'touch', 'chmod', 'chown', 'tee',
]);
const SINGLE_TARGET_FILE_COMMANDS = new Set([
'cat', 'head', 'tail', 'less', 'more', 'rm', 'touch', 'chmod', 'chown',
]);

function isLlmTrafficEvent(event: RuntimeAuditEvent): boolean {
return event.actionType === 'llm_request'
Expand Down
Loading
Loading