headroom icon indicating copy to clipboard operation
headroom copied to clipboard

feat(kompress): HEADROOM_KOMPRESS_BACKEND env + GPU/MPS auto-detect

Open mbachaud opened this issue 3 months ago • 2 comments

Summary

Closes #202.

Adds HEADROOM_KOMPRESS_BACKEND env var (auto / onnx / pytorch) and teaches auto mode to prefer the PyTorch backend when CUDA or Apple-Silicon MPS is available. Previously _load_kompress always tried ONNX first whenever onnxruntime was importable — which is always true for headroom-ai[proxy]. This left GPU-equipped users on a CPU-only path.

Selection order in auto (default):

  1. If PyTorch is installed AND (torch.cuda.is_available() OR Apple-Silicon MPS), prefer PyTorch; fall back to ONNX on failure.
  2. Else, prefer ONNX; fall back to PyTorch on failure.
  3. Raise ImportError if neither is available.

Apple Silicon detection uses platform.machine() == "arm64" and platform.system() == "Darwin", which Apple has committed to keeping stable across M-series generations (M1 / M2 / M3 / M4 / ...).

Invalid values (e.g. HEADROOM_KOMPRESS_BACKEND=tensorflow) log a warning rather than silently falling back to auto, so misconfiguration is visible.

Behavior on existing deployments

  • Linux/Windows CPU-only: no change — auto falls through to ONNX exactly like before.
  • NVIDIA + PyTorch installed: now auto-selects CUDA via PyTorch.
  • Apple Silicon + PyTorch installed: now auto-selects MPS.
  • Anyone can revert HEADROOM_KOMPRESS_BACKEND=onnx reverts.

Test plan

  • [x] Unit: HEADROOM_KOMPRESS_BACKEND=onnx forces ONNX
  • [x] Unit: HEADROOM_KOMPRESS_BACKEND=pytorch forces PyTorch
  • [x] Unit: auto + fake Apple-Silicon + MPS -> PyTorch
  • [x] Unit: auto + fake CUDA -> PyTorch
  • [x] Unit: auto + no accelerator -> ONNX (regression guard)
  • [x] Unit: auto + PyTorch load error -> falls back to ONNX
  • [x] Sanity: invalid env var logs warning and falls through
  • [x] Full kompress_compressor.py test file: 27/27 pass
  • [x] Full transforms/ test directory: 545 pass, 35 skip, 1 pre-existing fail unrelated to this PR (see below)
  • [ ] Manual: NVIDIA box — verify real CUDA path works end-to-end
  • [ ] Manual: Apple Silicon — verify real MPS path works end-to-end and confirm speedup from issue #202's field data (2206ms avg -> sub-1s expected)

I don't have Apple or NVIDIA CPU hardware locally. Requesting a maintainer or community reviewer to run the two manual checks before merge. Code-level correctness is fully covered by unit tests; wall-clock speedup is not.

Pre-existing CI state: tests/test_transforms/test_universal_json_crush.py::TestFullPipelineIntegration::test_number_array_via_compress fails on current `main` (verified via git diff origin/main — this PR does not touch universal_json_crush.py or its test file). The red CI on this PR matches what main already shows; not introduced here.

Docs

  • New `### Kompress backend selection` subsection in `wiki/configuration.md` covering the env var and the backend comparison table.
  • `CHANGELOG.md` entry under `[Unreleased]` / `Added`.

Commit structure

Five focused commits on the branch (keeping them separate for easier review + surgical revert if needed). Squash-on-merge is fine if that matches house style.

🤖 Generated with Claude Code

mbachaud avatar Apr 19 '26 06:04 mbachaud

Let me know if I need to make any changes.

Kompressor is a core part of Helix-Context stack, so having the auto-detect makes it "easier" on end users. Expecting to have a friend be able to test with DGX Sparks in a week or two to check off Nvidia devices

mbachaud avatar Apr 19 '26 20:04 mbachaud

Hi,

thanks for making changes - let me test it on my machine (I have a Mac) - and see if I can see the speedup.

chopratejas avatar Apr 20 '26 02:04 chopratejas