Skip to content
AyberkrkPublic

About

Explainable risk diagnostics for civil infrastructure anomaly detection + physics reasoning, trained on real FHWA bridge inspection data.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Cauren

Tests License: Apache 2.0

Cauren is a civil-engineering diagnostics and decision-support system. It ingests sensor/CBS ("bina/insaat" GIS) readings or bridge inspection records, normalizes them into a canonical feature schema, runs them through a generic anomaly-detection core, and hands the result to a domain physics/reasoning layer to produce an explainable, non-authoritative risk diagnosis.

Cauren does not issue final engineering decisions or automated enforcement. Diagnoses are meant for field inspection, permit review, and risk prioritization, never as an automated final engineering or enforcement decision.

raw sensor/CBS/inspection payload
        |
        v
  normalization  ->  agent router  ->  sector agent  ->  explainable diagnosis
                                            |
                                statistical anomaly core
                                            +
                                  domain physics layer

Two sector agents

cauren-civil cauren-bridge
Domain Buildings / construction Bridges
Required features 8 3
Live API today Yes (default) Registered, not exposed by default
Trained model No (statistical fallback only) Yes, see Evaluation
Ground truth used Narrow NYC CO administrative outcome in research only; no building-safety label Real 5-year deterioration outcome (FHWA)

cauren-bridge has a real, public, per-asset deterioration outcome from FHWA inspection history. The NYC research cohort adds an administrative CO outcome for new-building permits; it does not establish structural condition or safety and does not train the broader eight-feature building model.

cauren-civil feature schema (8 required)

Feature Unit What it represents
structural_risk_score ratio (0-1) Structural condition / risk level
inspection_finding_score ratio (0-1) Severity of open field-inspection findings
permit_status_score ratio (0-1) Permit/approval completeness
construction_progress_pct % (0-100) Construction completion level
infrastructure_connection_score ratio (0-1) Utility/infrastructure connection readiness
natural_hazard_score ratio (0-1) Seismic/flood/other natural hazard exposure
occupancy_safety_score ratio (0-1) Life-safety / occupancy risk
ground_stability_score ratio (0-1) Geotechnical/foundation stability

Optional: building_height_m, footprint_area_m2, soil_settlement_mm, utility_disruption_score.

cauren-bridge feature schema (3 required)

Feature Unit Source
structural_risk_score ratio (0-1) Derived from FHWA deck/superstructure/substructure condition ratings
ground_stability_score ratio (0-1) Scour risk/vulnerability derived from FHWA Item 113 (scour-criticality rating) via an explicit code -> risk mapping (categorical, not linear); 1.0 = poor stability / high vulnerability, same risk direction as cauren-civil's ground_stability_score
natural_hazard_score ratio (0-1) USGS seismic hazard (PGA) + FEMA flood zone

Evaluation

cauren-bridge's backbone (a masked-autoencoder encoder, a sequential 1-D convolution branch, and an RNN-Twin temporal-consistency branch, plus a small classification head) is trained and evaluated on a real dataset built from public sources, not synthetic or heuristic data.

Source FHWA National Bridge Inventory (2000-2023) + USGS seismic hazard + FEMA flood zone
States covered California, Iowa, Pennsylvania
Bridges 56,679, each with a real, independently-observed 5-year outcome
Supervised target deck_drop_5yr: did the deck condition rating actually drop within 5 years?
Train / validation / test 39,675 / 8,501 / 8,503 windows, split by bridge (no bridge appears in two splits)
Validation accuracy 77.3% (epoch 9, selected by validation supervised BCE -- see note)
Validation majority-class baseline 74.1%
Test accuracy 77.3% (untouched test split, never used for training or checkpoint selection)
Test majority-class baseline 74.0%
Test precision / recall / F1 74.6% / 19.2% / 30.5% (prediction threshold 0.5)
Test balanced accuracy 58.4%
Test confusion matrix TP 423, FP 144, TN 6151, FN 1785
Parameters 176,196

77.3% against a ~74% baseline is a modest, genuine improvement, not a spectacular one, and that is the point: the input window for each prediction is strictly limited to years at or before the outcome's reference year, so the model cannot see the future it is predicting. A model claiming much higher accuracy on this kind of task without a similar leakage check should be treated with suspicion. Validation and test accuracy landing within a hundredth of a point of each other is a good sign that the checkpoint wasn't overfit to whichever split selected it. See data/cauren_bridge/dataset_summary.json (built locally, gitignored because of size, see Rebuilding the bridge dataset) and cauren_core/checkpoints/cauren_bridge_backbone_bundle.pt for the trained artifact; tools/evaluate_cauren_bridge_checkpoint.py reproduces the test-split numbers above from that checkpoint.

Accuracy alone overstates this model. deck_drop_5yr is imbalanced (~25% positive: most bridges don't deteriorate within 5 years), so a model that always predicts "no drop" already scores ~74-75% accuracy without ever finding a real deterioration case. The 77.3% headline number is only 3 points above that trivial baseline, and the confusion matrix above shows why: at the 0.5 prediction threshold, the model finds 423 of the 2,208 real positive cases in the test split and misses 1,785 of them (19.2% recall), while balanced accuracy (58.4%, the average of recall and specificity) is much closer to the 50% chance line than the 77.3% raw-accuracy figure suggests. Precision (74.6%) is respectable when the model does flag a bridge, but it flags relatively few. This is a real limitation of the current checkpoint, not a reporting error, and it's why tools/evaluate_cauren_bridge_checkpoint.py reports precision, recall, F1, balanced accuracy, and the raw confusion matrix alongside accuracy rather than accuracy alone.

Checkpoint selection. For a dataset without a real supervised target (cauren-civil), the best checkpoint is the epoch with the lowest validation reconstruction (MAE) loss. For cauren-bridge, which has one (deck_drop_5yr), the best checkpoint is instead the epoch with the lowest validation supervised loss (binary cross-entropy on deck_drop_5yr), not reconstruction loss -- the two can diverge (an epoch can reconstruct the input window better while forecasting the outcome worse), and the model is explicitly evaluated on the forecasting task. BCE is used over raw accuracy for selection because the target is imbalanced (~25% positive); reconstruction loss is still recorded every epoch for diagnostics (cauren_core/training.py's train_summary["validation_loss_history"]), and the selected metric and its value/epoch are recorded in train_summary["selection_metric"] / train_summary["best_selection_metric_value"]. The figures above were measured after both this rule and the Item 113 scour-direction fix (see ground_stability_score above) landed, by rebuilding the dataset from the original FHWA/USGS/FEMA sources and retraining from scratch:

python3 tools/train_cauren_core.py \
    --dataset-dir data/cauren_bridge --agents cauren-bridge \
    --output-path cauren_core/checkpoints/cauren_bridge_backbone_bundle.pt \
    --epochs 20 --batch-size 64 --learning-rate 0.0004 \
    --early-stopping-patience 5 --min-epochs 5
python3 tools/evaluate_cauren_bridge_checkpoint.py

(the retraining accuracy landed within 0.02 points of the pre-fix number measured under the old reconstruction-loss selection rule and the old, direction-inverted scour feature -- on this dataset the two fixes turned out not to move the headline accuracy much, which is itself a useful, honestly-reported data point, not a reason to have skipped verifying it).

Status and honest limitations

This project is a working research prototype, not a finished product.

  • cauren-civil has no pretrained model yet. cauren_core/checkpoints/ ships without one for it. Out of the box, the anomaly-detection core runs on its statistical fallback (median/spread-based calibration and pattern-shape scoring), not a trained neural backbone.
  • A single-snapshot request has no time series to analyze. If you send one reading per feature with no history, the core reports insufficient_temporal_history / snapshot_only instead of a fabricated pattern verdict, and the risk score is driven mainly by the physics layer's read of the values you sent. Send a real time series (seq_len distinct timestamps) for genuine pattern-shape detection (drift, spike, oscillation, flatline, and so on).
  • LLM/ (rocket_llm) is pre-alpha. No CLI is exposed yet -- pyproject.toml deliberately has no [project.scripts] section, commented out until each backing module exists (see the file for the planned list). Only LLM/src/rocket_llm/pipelines/training.py is implemented so far.
  • cauren-civil does not derive its 8 scores from raw sensor/CBS data, it consumes them. The required features are expected to already exist (from a field inspection, an engineering assessment, another model) before they reach Cauren. This was validated end to end against a real public dataset (US Census building-permit records via Hugging Face, thanna94/us-building-permits): it cleanly produced two of the eight, permit_status_score and construction_progress_pct, but the other six have no off-the-shelf public source, and the pipeline correctly reported them as missing_features instead of inventing values. What Cauren adds on top of scores you already have is anomaly-pattern detection over time, physics-relation reasoning between scores, and an explainable diagnosis, not score generation from raw measurements.

Bridge model comparison

tools/benchmark_cauren_bridge_models.py compares the bridge models on the same bridge-disjoint splits using PR-AUC, Brier score, calibration error, and recall and false alarms when reviewing the highest-risk 5% or 10% of bridges. On the local v3 split, the saved hybrid scored PR-AUC 0.532, Brier 0.1640, and ECE 0.0119. Same-input histogram boosting scored PR-AUC 0.5506 and Brier 0.1609. A paired bridge-cluster bootstrap puts the hybrid-minus-booster PR-AUC difference at -0.0184 (95% interval -0.0263 to -0.0096) and Brier difference at +0.00315 (0.00186 to 0.00433), favoring boosting on this split; the hybrid has lower ECE. Strict leave-one-state-out refits showed uneven transfer, so no new probability is exposed through the API. Protocol, state results, uncertainty intervals, and reproduction steps are in docs/bridge-model-benchmark.md.

NYC DOB research dataset

The new-building permit cohort links BIS and DOB NOW permits to both certificate feeds, with violation and complaint records restricted to the cohort's BINs. The 2026-09-29 snapshot contains 94,910 buildings. In a five-year horizon, 14,612 have a qualifying non-temporary CO record and 73,336 have mature follow-up without one; 6,962 remain censored or ambiguous. This is an administrative outcome, not a safety label. The 57-record linkage review, historical DOB field versions, and FEMA map vintage checks remain open, so the dataset is not ready for model training. See the NYC dataset method and audit and the combined research plan.

What ships with every diagnosis

Beyond the risk score itself, each diagnosis carries the checks that were run on the way to it, so a reviewer can see how much to trust it:

Output key What it answers
quality_control Were the inputs usable? Required-feature completeness, unit-range sanity, flatline/stuck sensors, how many sensor names failed to normalize, how many raw readings were rejected.
uncertainty_report How much should the score be trusted? A confidence band around the risk score derived from data quality, the core's own confidence, and how much of the physics schema had evidence.
physics_evidence Why is the score what it is? The domain relations that fired, and which required features were missing.

Structural-health inputs are optional and separate. cauren_physics/oma.py identifies natural frequencies from one vibration channel and estimates damping from a Welch spectrum. Damping is reported only when the measured half-power bandwidth is resolved. With multiple aligned channels, optional Timoshenko 2.x support adds Frequency Domain Decomposition (FDD) mode shapes. cauren_physics/fe_reference_model.py compares measured modes with a shear-building model and reports a separate model-consistency ratio. FE frequencies do not replace a measured drift baseline. Measured baseline drift is a decision-support signal for human review, not a damage verdict.

Cauren can use Timoshenko for shared structural calculations: single-channel modal identification, mode pairing, multi-channel FDD, shear-building frequencies, and the global stiffness update summary. Cauren-specific risk relations and model interpretation remain in Cauren. Timoshenko is optional; the Cauren package and single-channel fallback work without it. Install Timoshenko Engine into the same Python environment to enable the shared engine path:

python -m pip install timoshenko-engine

The distribution name is timoshenko-engine, while the import name is timoshenko.

# Full explainable review over real data (repo ships a small public dataset)
python3 tools/run_explainable_review.py \
  --wide-csv data/public_sources/normalized/civil_public_core_assets.csv

# ...or from a raw sensor feed, with a modal-drift check folded in
python3 tools/run_explainable_review.py \
  --sensors-json my_sensors.json \
  --vibration-csv my_vibration.csv --sampling-hz 100 \
  --baseline-frequencies-hz 2.01

# Multi-channel FDD uses columns in channel-major order and requires Timoshenko 2.0+
python3 tools/run_oma_identification.py \
  --input-csv my_channels.csv --columns north east west --sampling-hz 100

# FE consistency is reported separately from the measured drift baseline
python3 tools/run_oma_identification.py \
  --input-csv my_vibration.csv --column acceleration --sampling-hz 100 \
  --baseline-frequencies-hz 2.01 \
  --fe-story-masses-kg 50000 --fe-story-stiffness-n-per-m 8000000

These are decision-support signals for routing an asset to a human reviewer, not damage verdicts.

Repository layout

Path What it is
cauren_core/ Sector-agnostic runtime: contracts, orchestrator, statistical anomaly core, the trainable backbone, normalization
cauren_agents/ Sector agents (civil, bridge) and the router
cauren_physics/ Domain physics/reasoning layer, one module per sector
api/ FastAPI service exposing the pipeline over HTTP
tools/ Dataset build and training scripts
data/ Dataset manifests, sources, and (locally built) training data
tests/ Test suite (133 tests with the minimal dependency set; 152 total once pandas/torch/PyYAML are also installed)
operations/ Architecture and data-contract reference docs
LLM/ Pre-alpha advisory-LLM sub-project, not yet functional

API endpoints

Method Path Purpose
GET /health/liveness, /health/readiness Health checks
GET /agents List registered sector agents
POST /calibrate Normalize/calibrate a sensor payload without running physics
POST /diagnose Full pipeline: route, calibrate, classify, physics, compose. Accepts either a sensors list or a CBS-shaped building record (building_id + risk_assessments/project_permit/etc, converted internally)

api/app.py is a thin, unauthenticated wrapper around cauren_core.CaurenPipeline -- there is no built-in login, token, or per-tenant access control. It's meant to be run locally or behind whatever the operator puts in front of it (a reverse proxy, an auth layer), not exposed to the open internet as-is.

Run the API with Docker

For a local trial, Docker Compose builds the API image and binds it to 127.0.0.1 so it is not published on your network by default:

docker compose up --build

When the service is ready, open http://127.0.0.1:8000/docs or check http://127.0.0.1:8000/health/readiness. Stop it with:

docker compose down

This local setup does not add authentication or tenant isolation. Do not change the port binding or expose the service publicly without placing appropriate access controls in front of it.

Quick start

The project is pip-installable from the repo root (pyproject.toml declares cauren_core, cauren_agents, cauren_physics, api, and tools as packages, with no required dependencies of its own):

pip install -e '.[test]'
python3 -m pytest tests/            # 133 tests, no network/GPU/numpy/torch required (152 total with pandas/torch/PyYAML also installed)

pip install -e '.[api]'             # only needed to actually run the API
python3 api/app.py                  # or: uvicorn api.app:app --reload

Installing individual packages directly still works if you'd rather not use the extras above:

pip install fastapi pydantic pytest httpx
pip install uvicorn                 # only needed to actually run the API

numpy is optional: cauren_physics/oma.py uses it for faster FFTs when present, and falls back to a pure-Python DFT otherwise. Nothing else in the pipeline touches it. Building/training against the bridge dataset needs the dataset extra (pip install -e '.[dataset]', i.e. numpy/pandas/scipy), and training the core backbone needs the backbone extra (torch) -- see the next section.

Rebuilding the bridge dataset

data/cauren_bridge/ is not checked in (it is ~1.5GB and fully reproducible from public sources). To rebuild it you need pandas and scipy in addition to the packages above (pip install -e '.[dataset]'), and to pull the source data yourself:

  • FHWA National Bridge Inventory: sweetapricity/bridgedeck-nbi on Hugging Face
  • USGS seismic hazard: sweetapricity/bridgedeck-nshm
  • FEMA flood zone: sweetapricity/bridgedeck-nfhl

See data/cauren_bridge/dataset_summary.json's source field and cauren_physics/bridge.py's module docstring for the feature derivations, and tools/train_cauren_core.py for the training entry point once the dataset exists locally.

Contributing

See CONTRIBUTING.md.

License

Apache 2.0. See LICENSE.

About

Explainable risk diagnostics for civil infrastructure anomaly detection + physics reasoning, trained on real FHWA bridge inspection data.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages