I build verifiable training loops: pipelines that turn a single historical source — an 1890 Dakota grammar, an 1865 Cree dictionary, a 1959 railroad rulebook — into extracted rules, labeled tasks, deterministic reward functions, and published models. Plus the project's other half: Math-To-Manim ( 2,400+), a prompt-to-animation engine that turns questions into computed mathematical film.
| 703 public repos |
2,400+ stars on Math-To-Manim |
17 / 8 / 7 HF models, datasets, Spaces |
82M+ tokens through one GRPO run |
3 languages brought to RL |
What I do that employers actually hire for:
- Model training & fine-tuning — GRPO / RL post-training with deterministic, decomposed reward functions; LoRA adapters from 0.5B to 35B on Tinker and Prime Intellect; full run cards, reward ledgers, and audits published on Hugging Face and W&B.
- Data labeling & dataset engineering — VLM extraction from archival scans, orthography-preserving labeling, synthetic Q&A expansion, structural holdouts, hash-addressed dataset artifacts with citations intact.
- Reward/verifier design — grammar rules compiled into executable, per-component reward channels (no LLM judge; every gradient is inspectable).
- Visualization that explains — Manim render pipelines, RL training-curve dashboards, LiDAR terrain viewers.
Ask a question → get a freakin' movie. Nothing keyframed — every frame is integrated, simulated, or derived. 2,400+ stars.
The signature move: grammar as a reward function. Take a source document, compile its rules into verifiable tasks, and let GRPO optimize against per-component reward channels — orthography, morphology, semantics — with zero LLM-judge fuzz. The reward ledger reconciles to the step.
reward = (
0.4 * character_preservation + # orthography: ŋ š ć ḣ preserved?
0.4 * affix_accuracy + # morphology: correct affixes applied?
0.2 * semantic_correctness # semantics: meaning vs. ground truth
) * difficulty_multiplier # curriculum weight, 1.0x → 2.0x| Model | Params | Method | Verified result |
|---|---|---|---|
| Laguna-XS.2-Adaption-Dakota-QA-GRPO | XS | GRPO, Prime Hosted Training | Reward 0.283 → 0.433, char-F1 0.327 → 0.635 |
| Qwen3.6-35B-A3B-Dakota1890-GRPO | 35B | GRPO, Tinker | 82.05M tokens, audited reward channels |
| Cree1865 | 30B-A3B | Modified GRPO, Tinker | 800-step synthetic-expansion run, live W&B |
| Qwen3-4B-RailRoadEngineer1959 | 4B | LoRA, volume2gym lineage | Rulebook-compiled task families |
| Qwen3-0.6B-Dakota-Grammar-RL-400 | 0.6B | GRPO, Prime Intellect | 400 steps, +150% reward, 97.9% morphology accuracy |
| nanochat-AquaRat | nano | RL, AQuA-RAT | GSM8K-style → multiple-choice algebra |
volume2gym is the general compiler behind the language work: any structured volume becomes an RL gym. Sections name the world, rules constrain action, procedures encode order, exceptions define edge cases. The compiler emits cited knowledge units, six task families, grouped holdouts, deterministic reward ledgers, and SFT/GRPO trainer exports — hash-addressed and tamper-evident.
| Task family the compiler emits | What it tests |
|---|---|
standard_operation |
Correct ordinary application |
edge_case |
Boundary conditions and missing facts |
conflict_resolution |
Compatible resolution of constraints |
exception_handling |
Exception triggers vs. normal boundaries |
violation_check |
Missing requirements, forbidden actions, bad order |
adversarial_distractor |
Rejection of plausible but unsupported instructions |
The 1959 Consolidated Code of Operating Rules lineage: 536 extracted rules → 2,708 scenarios → gym → Qwen3-4B adapter → Rule 99 contract fixture on Hugging Face.
| Dataset | What it is | Shape |
|---|---|---|
| adaption-dakota-english-qa | Remastered Dakota–English QA for instruction tuning & GRPO | 1,953 examples |
| dakota-bilingual-qa | Bilingual QA pairs from the 1890 dictionary | 2,445 examples, train/val |
| Stoney10kRL | 2026 Stoney Nakoda RL fine-tuning package | 8,000 train / 2,000 val |
| StoneyNakoda45k | Community-in-the-loop language dataset | 25–50K size class |
| volume2gym-railroad-1959 | Rule 99 artifact-contract fixture with ledgers | 6 train / 1 held-out |
| synthetic_stoney_data | Synthetic Q&A bootstrap resource | JSONL |
A research line in three languages: give an endangered language one good historical book, and train. VLM extraction reads the scans — diacritics, letterpress ligatures, and all — then synthetic Q&A multiplies the surface area, then GRPO with a rubric built from the book itself. The final stage belongs to the community: speakers correct the model, and the corrections become the next training round. The model is a toddler that has read the book cover to cover; the community teaches it the rest.
| Project | Source volume | Public artifacts |
|---|---|---|
| Dakota1890 | Riggs 1890 Grammar & Dictionary of the Dakota Language | 35B adapter · Laguna run card |
| Cree1865 | Watkins 1865 Dictionary of the Cree Language | HF model · W&B run · explained dashboard · inference Space |
| StoneyNakoda | Contemporary speakers + historical survey material | Stoney10kRL · StoneyApp |
| Railroad Engineer 1959 | 1959 Consolidated Code of Operating Rules | Qwen3-4B LoRA · dataset fixture |
Handwriting and OCR lineage runs through the repo list too — PyLaia (handwritten document analysis), deepseek-ocr, olmocr (PDF linearization for training data), and a reproduction of LeCun 1989 handwritten zip-code recognition — the ancestor of all of this.
| Project | What it shows |
|---|---|
| lidar2 | Map-driven LiDAR visualizer — OpenTopography DEM → multi-layer 3D terrain point clouds (React, Three.js, custom GLSL elevation shaders) |
| maplibre-gl-lidar | MapLibre plugin for visualizing LiDAR point clouds |
| openArchive | Research UX over BC & Alberta archive collections |
| AlbertaWorkspaceAgent | Agent-native workspace experiments for Alberta research workflows |
| Live demos (Spaces) | Try it |
|---|---|
| StoneyApp | Stoney Nakoda community-in-the-loop app |
| Cree1865-Tinker-Inference | Sample from the Cree1865 training run |
| Dakota-.6B | Dakota grammar RL demo |
| AskAboutCIL | Community-in-the-loop method explainer |
Every training run ships with its curves public — reward channels, entropy, ledger audits, per-step components.
Live — refreshed every 6 hours by a GitHub Action from CNBC, Reuters, and FT feeds.
| Category | Date | Headline |
|---|---|---|
| Market | Jul 23, 2026 | Odds of Federal Reserve rate hike surge as oil prices rip higher |
| Market | Jul 23, 2026 | JPMorgan report finds dramatic jump in AI-themed ETFs — despite rough quarter |
| Market | Jul 23, 2026 | Shortsighted stock market can no longer brush off war: 'It's too hard to ignore $100 oil' |
| Market | Jul 22, 2026 | John Paulson says we are in the early stages of a long-term bull market for gold |
| Market | Jul 22, 2026 | Kalshi launches election hub for prediction markets ahead of midterms |
| Finance | Jul 23, 2026 | Oil hits $100 for first time since May while US stocks slide |
| Finance | Jul 23, 2026 | Houthi attacks threaten Saudi Arabia’s oil lifeline |
| Finance | Jul 23, 2026 | US oil refineries run at breakneck speeds as wars choke fuel supplies |
| Finance | Jul 23, 2026 | Japan awakes |
| Finance | Jul 23, 2026 | Is Trump winning the tariff wars? |
GitHub · Hugging Face · Weights & Biases · LinkedIn · X · Kaggle
Open source is the portfolio. The best entry points are the project READMEs, model cards, dataset cards, W&B runs, and demos linked above.



















