Skip to content

feat: add low-VRAM Qwen, Gemma, and Llama profiles#4

Merged
waqasm86 merged 2 commits into
mainfrom
codex/multi-model-gguf-20260719
Jul 19, 2026
Merged

feat: add low-VRAM Qwen, Gemma, and Llama profiles#4
waqasm86 merged 2 commits into
mainfrom
codex/multi-model-gguf-20260719

Conversation

@waqasm86

Copy link
Copy Markdown
Collaborator

Adds Modelfile.llama3.2-1b, dedicated Gemma and Llama overlays, safe registry-backed Modelfile rendering, multi-model documentation, and Helm tests. The validated profiles keep observed total NVIDIA use below the 900 MiB guardrail by offloading remaining layers to CPU/system RAM.\n\nValidation: Go tests and vet, Helm lint, three profile renders, JSON validation, and live sequential deployment passed.

@waqasm86
waqasm86 merged commit 315ea42 into main Jul 19, 2026
2 checks passed
@waqasm86
waqasm86 deleted the codex/multi-model-gguf-20260719 branch July 19, 2026 07:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant