Why Qwen 3.5 and not 3.6?
Timon
KeyboardMasher
AI & ML interests
None yet
Recent Activity
new activity about 11 hours ago
prism-ml/Ternary-Bonsai-27B-gguf:Update metadata (tokenizer.ggml.model) new activity about 12 hours ago
prism-ml/Bonsai-8B-gguf:Poor performance new activity 4 days ago
LiquidAI/LFM2.5-VL-3B-GGUF:Temperature is last for Llama.cppOrganizations
None yet
Update metadata (tokenizer.ggml.model)
1
#50 opened about 2 months ago
by
Lethaulte
Poor performance
8
#6 opened 3 months ago
by
KirMas
Temperature is last for Llama.cpp
#4 opened 4 days ago
by
KeyboardMasher
Correct Samplers Order?
#64 opened 6 days ago
by
KeyboardMasher
mmproj files are missing
#1 opened 25 days ago
by
KeyboardMasher
version for 16 GB VRAM
1
#5 opened about 1 month ago
by
KeyboardMasher
Feedback
1
#1 opened about 1 month ago
by
KeyboardMasher
token_embd.weight is quantized too much
#2 opened 2 months ago
by
KeyboardMasher
commented on Gemma 4 VLA Demo on Jetson Orin Nano Super 5 months ago
At 4 bit or lower use IQ-type quant. The math is more advanced and quantization error is lower. You can double the context with -ctk q8_0 -ctv q8_0for virtually no loss of quality and speed.
Gemma 4 seems to work best with high temperature for coding
๐ 1
8
#21 opened 6 months ago
by
Reverger
Recommended sampler?
4
#4 opened 6 months ago
by
mratsim
Older quants get in the way
2
#1 opened 7 months ago
by
KeyboardMasher
Error with built-in Web UI
2
#3 opened about 1 year ago
by
KeyboardMasher
Thanks for IQ4_NL
โค๏ธ 1
#1 opened about 1 year ago
by
KeyboardMasher
128k Context GGUF, please?
4
#2 opened over 1 year ago
by
MikeNate
Update README.md
#1 opened over 1 year ago
by
KeyboardMasher
Other Imatrix quants (IQ3_XS) ?
๐ 3
6
#1 opened over 1 year ago
by deleted
reacted to bartowski's post with ๐ over 1 year ago
Post
40042
Access requests enabled for latest GLM models
While a fix is being implemented (https://github.com/ggml-org/llama.cpp/pull/12957) I want to leave the models up for visibility and continued discussion, but want to prevent accidental downloads of known broken models (even though there are settings that could fix it at runtime for now)
With this goal, I've enabled access requests. I don't really want your data, so I'm sorry that I don't think there's a way around that? But that's what I'm gonna do for now, and I'll remove the gate when a fix is up and verified and I have a chance to re-convert and quantize!
Hope you don't mind in the mean time :D
While a fix is being implemented (https://github.com/ggml-org/llama.cpp/pull/12957) I want to leave the models up for visibility and continued discussion, but want to prevent accidental downloads of known broken models (even though there are settings that could fix it at runtime for now)
With this goal, I've enabled access requests. I don't really want your data, so I'm sorry that I don't think there's a way around that? But that's what I'm gonna do for now, and I'll remove the gate when a fix is up and verified and I have a chance to re-convert and quantize!
Hope you don't mind in the mean time :D