Skip to content
View achirothmane's full-sized avatar
🤗
🤗

Block or report achirothmane

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
achirothmane/README.md
Othmane Achir — Evidence-Gated Systems

Verification & control infrastructure for AI, automation, and high-consequence software

I build systems that decide when there is enough evidence to act — and when the correct answer is UNKNOWN.

CI Retry Gate Agent Action Guard PostgreSQL Change Safety


What I build

I build verification and control systems for software that should not act on weak evidence.

My current work focuses on:

AI-agent control · CI reliability · causal verification · developer tooling · high-consequence automation

I work primarily with:

Python · GitHub Actions · PostgreSQL · CI/CD · AI/LLM systems

The recurring engineering question behind my projects is:

What evidence must be established before this system is allowed to act?

Open to

Verification / reliability / developer-infrastructure work, applied AI systems, and selected technical collaborations where evidence, control, and auditability matter.


Most relevant systems

Problem
CI systems often retry failures without knowing whether retrying is actually justified.

Built
An evidence-gated decision layer that checks provenance, causal evidence, side effects, and retry limits before granting rerun authority.

Why it matters
Fewer blind reruns, clearer failure handling, and auditable retry decisions.

failure → evidence → decision → retry / block

Problem
An AI agent may reach a real-world consequence through a path that bypasses the approval boundary intended to control it.

Built
A scanner and runtime witness for modeled consequence paths, expected boundaries, bypasses, and unresolved evidence.

Why it matters
Authorization controls are only useful if alternate paths cannot silently route around them.

agent → path analysis → boundary → counterexample / covered / unknown

Problem
A workload regression after a PostgreSQL change is easy to observe and easy to misattribute.

Built
Comparable workload windows, stable SQL fingerprints, controlled causal experiments, plan-variant analysis, and explicit confounder handling.

Why it matters
The tool separates “a regression happened” from “we have enough evidence to say why.”

before/after → regression → causal isolation → evidence strength

Problem
A legal-AI answer can still look plausible while its citations, authority strength, treatment, or claim support become weaker.

Built
Differential regression testing across baseline and candidate outputs, with explicit UNKNOWN and WORLD_CHANGE states.

Why it matters
A citation existing is not the same as the citation being sufficient support for the claim.

baseline → candidate → authority diff → block / unchanged / world change

Also building

Private Code Modernization Factory — evidence-first modernization analysis with constrained patch proposals, differential verification, and escalation when automation has not earned authority.


Technical focus

Python GitHub Actions PostgreSQL CI/CD AI Agents LLM Evaluation Verification Developer Infrastructure


The engineering thesis

A recurring failure pattern appears across AI agents, CI systems, databases, legal AI, and software automation:

Plausibility is not authority. A system should act only when the evidence required for that action has actually been established.

The architecture I keep returning to is:

Input / Event / Proposed Action
              │
              ▼
      Evidence Collection
              │
              ▼
         Verification
        ┌─────┴─────┐
        │           │
   sufficient   insufficient
        │           │
        ▼           ▼
Authority/Policy   UNKNOWN
        │        BLOCK / ESCALATE
        ▼
    Decision Gate
     ┌───┼───┐
     ▼   ▼   ▼
   ALLOW BLOCK HUMAN
        │
        ▼
     Execution
        │
        ▼
 Outcome Verification
        │
        ▼
   Audit / Replay

Design rules I care about

Explicit uncertainty

UNKNOWN is a valid engineering outcome. It is safer than inventing certainty from weak evidence.
Fail closed

High-consequence actions do not silently inherit permission when evidence is incomplete.
Read-only first

Observe and prove value before enabling mutation, reruns, quarantine, merge, deploy, or other write authority.
Differential verification

Prefer controlled before/after evidence over plausible explanations.
Causal isolation

When attribution matters, test competing explanations instead of treating correlation as cause.
Audit & replay

Important decisions should retain enough evidence to inspect and reproduce how authority was granted.

Connect

LinkedIn

Open to selected technical collaborations in verification, reliability, developer infrastructure, and applied AI systems.

---

Evidence before action.

Build → test → falsify → strengthen the evidence → automate only what has earned authority.

Independent tools first. Shared platform only when repeated real-world usage proves the same primitives belong together.

Pinned Loading

  1. workflow-failure-lab workflow-failure-lab Public

    GitHub Actions retry safety + flaky test intelligence: safe reruns, JUnit evidence, quarantine lifecycle, CODEOWNERS routing, and deduplicated issue tracking.

    Python 3