AI-generated code is easy to produce and expensive to trust.
Shipispec builds reliability tooling for AI-assisted software engineering: deterministic recovery where possible, model-based repair where necessary, and reproducible benchmarks for knowing the difference.
| Project | Problem | Evidence |
|---|---|---|
| tsfix | Repairs the TypeScript failures that code-generation agents leave behind, with knowledge of the libraries actually installed | Real-world failure benchmark, deterministic + LLM layers, CI integration, security policy, changelog |
generator → compile / test failure → deterministic quick-fix (language server)
→ contextual LLM repair (installed-library versions injected)
→ verification (re-compile, re-test)
Deterministic first, stochastic second, verify always. The LLM layer is opt-in and metered.
| Workload | Pass rate | Notes |
|---|---|---|
| Single-file repairs | 98.6% | the supported path |
| Multi-file repairs | 40.0% | the current limitation, reported on purpose |
| Aggregate suite | 81.4% | real failure fixtures, n=3 per cell |
Cost: under $0.005 per single-file repair. Fixtures, runner, model configuration and raw results live in
tsfix/benchmark.
Sibling organization: Shipi18n applies the same discipline — deterministic invariants, a separately benchmarked LLM layer, validation against real repositories — to translation files.
Apache-2.0 · maintained by owgreen-dev