Edward Paolo Guevarra PRO
PaoAI
AI & ML interests
None yet
Recent Activity
updated a model 1 day ago
PaoAI/Qwen3.8-27B-PaoAI-ROCmFP4-STRIX-BALANCED-GGUF posted an update 4 days ago
New on the Hub: **Qwen3.8-Flash-Next STRIX BALANCED-2.1** for AMD Strix Halo (Ryzen AI Max+ 395, 128 GB).
It's the same BALANCED-2 recipe with its always-on dense weights stored in 8-bit: 78.3 GB instead of 80.8, the same perplexity (−0.16 %, within error), and faster writing. It runs on a new ROCm/HIP engine built on @ilintar's Strix Halo llama.cpp branch plus one small fix of ours for the MTP check step.
Measured on one box with one fixed method: the full 262,144-token window checked at every depth (8K → 256K), a needle found at 260K, a 3-run coding exam graded by running the code (median 100/100), and a 55-task quality bench. Writing runs 17.5–46.7 t/s depending on depth and answer kind; reading runs 301–870 t/s.
👉 https://huggingface.co/PaoAI/Qwen3.8-Flash-Next-PaoAI-STRIX-BALANCED-2-GGUF
🛠 Engine: https://github.com/guevae2/paoai-qwen38fn-rocm-engine
Big thanks to @ilintar (Piotr Wilkin) for the strix-halo branch and ROCm runtime. It's his engine, and our part is one 3-line fix. Thanks also to @unsloth for the MTP draft sidecar and to Halogen for the 8-bit-dense idea. updated a model 4 days ago
PaoAI/Qwen3.8-Flash-Next-PaoAI-STRIX-BALANCED-2-GGUFOrganizations
None yet