Skip to content

Instantly share code, notes, and snippets.

@deepu105
deepu105 / strix-halo-halogen-gufo-llamastash.md
Last active October 6, 2026 20:03
Handover: Halogen + gufo behind LlamaStash on Strix Halo (Qwen3.8 Flash-Next and 27B)

Handover: Halogen + gufo behind LlamaStash on Strix Halo

You are setting up a local LLM stack on an AMD Strix Halo machine (Ryzen AI Max+ 395, Radeon 8060S, gfx1151, 128 GB unified memory) running native Linux. When you finish, the user can start any of these from LlamaStash (TUI, CLI, or its OpenAI/Anthropic proxy), each with ready-made presets:

LlamaStash row Engine Weights
flash-next-halogen Halogen (Docker image) Halogen native v2 .hgn (Qwen3.8-Flash-Next)
Qwen3.8-Flash-Next-UD-Q4_K_XL gufo (native build) Unsloth UD-Q4_K_XL GGUF + Unsloth shared MTP head
Qwen3.8-27B-UD-Q6_K gufo (native build) Unsloth UD-Q6_K GGUF + z-lab DFlash2 draft