Use llama.cpp profiling and numerical evidence to guide compute-unit boundaries, shared resources and host/device work. Software also owns operation mapping, backend/runtime and host integration.
- Read the current assignment and artifact locations.
- Work on a branch and open a PR for
@abhinavnandwaniusing CONTRIBUTING.md. Main requires a code-owner approval; admins can bypass.
| Location | Purpose |
|---|---|
| experiments/llama-cpp/ | Recorded baseline in 2026-09-22/. Put each new run in a separate dated or named directory with a manifest, commands, source/model hashes, host/backend, inputs, results and limitations. |
| research/compute-mapping/ | Operation/dependency maps, offload comparisons and recommendations. Link each claim to a run or source trace; separate measured time/traffic from estimates. |
A reproducible CPU/Metal profiling package and the small MAC reference example exist. CPU time shares are not accelerator area or speedup. Mixed Q4_K_M results do not validate the slide's custom INT4/group-64 proposal.
Shared diagram · Software evidence
CHARTER.md and OBJECTIVES.md describe the longer-term purpose. Current issues and the starting guide specify the work assigned now. SETUP.md describes existing example commands and scope.