Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
jaithree
jaithree
2
3
24
Follow
AuraZool's profile picture
PhysiQuanty's profile picture
Quazim0t0's profile picture
4 followers
·
26 following
AI & ML interests
"In the beginning was the word. Then came the **** word processor. Then came the thought processor. Then came the death of literature. And so it goes" -- Hyperion - Dan Simmons
Recent Activity
reacted
to
SeaWolf-AI
's
post
with 🔥
about 1 hour ago
POCKET now speaks Gemma 4 — a 26B model that loads in every app, and runs on your PC with no GPU We're adding a Gemma-4 sibling to POCKET: POCKET-26B, built from Google's Gemma-4-26B-A4B (Apache-2.0). Our flagship POCKET-35B is a Qwen-family MoE and needs a recent llama.cpp; POCKET-26B trades a little size for the thing people kept asking for — it just loads, everywhere, today: Ollama, LM Studio, PocketPal, MLX, any stock llama.cpp. No fork, no bleeding-edge runtime, no CUDA, no cloud. It's a sparse Mixture-of-Experts (25.2B total, ~4B active per token), so the work per token stays small — a real 26B that generates on a CPU with no graphics card. Two things make it stand out: 1) Universal compatibility. Gemma 4 is a standard, widely-supported architecture, so POCKET-26B runs on the tools you already have — no waiting for your app to add a new model type. 2) Quality that survives compression. Measured GPQA-Diamond (198 q, greedy): • Full base: 67.7% • POCKET-26B Q4_K_M (17 GB): 67.7% — lossless • POCKET-26B Q2_K (11 GB): 67.2% — near-lossless, at 11 GB Live, on a CPU-only box (our demo Space — POCKET-26B vs Bonsai-27B, same machine, same stock llama.cpp): POCKET-26B ≈ 19 tok/s vs Bonsai ≈ 6 tok/s → about 3× faster generation, no GPU. (Honest notes: shared CPU box, sequential race; a dedicated machine is faster.) Where it fits in the family: • POCKET-35B (Qwen MoE) — bigger, top-tier, needs a recent llama.cpp. • POCKET-26B (Gemma 4) — loads in any app, quality-robust when compressed. The demo runs the Q4_K_M build; Q2_K (11 GB) is the smallest footprint. For a true ≤8 GB phone, the 5 GB POCKET-KR (Qwen) is still the pick. Try it and grab it: 🖥️ Live demo (Gemma4-based, answering on a CPU, no GPU): https://huggingface.co/spaces/FINAL-Bench/POCKET-26B-CPU 📦 POCKET-26B-GGUF (Q4_K_M 17 GB · Q2_K 11 GB): https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF 📚 POCKET collection: https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6
reacted
to
badaoui
's
post
with 👍
about 6 hours ago
432 GB of ultra-fast HBM4 and up to 23.3 TB/s of memory bandwidth on a single GPU 🤯. Two weeks ago, we got early access to AMD's new Instinct MI455X, and our first goal was simple: make sure 🤗 Transformers works on day one. Over the past few weeks, we worked closely with the AMD team to validate the platform, enable Flash Attention, add torchcodec support for multimodal models, and resolve issues uncovered during testing. The result: ✅ 99.5% success rate across our 24 core Transformers model architectures - already on par with our daily CI on previous AMD and NVIDIA platforms. The hardware is just as exciting. With 432 GB of HBM per GPU, our early capacity experiments showed more than 3× the concurrent long-context requests compared to MI300, thanks to the much larger KV cache capacity. A huge thanks to the AMD team for the early access and the great collaboration! Read the full blog 👇 https://huggingface.co/blog/badaoui/transformers-on-amd-mi455
liked
a model
about 15 hours ago
ReadyArt/gemma-4-31B-it-scotoma-GGUF
View all activity
Organizations
None yet
models
1
jaithree/grug-9b-mlx-q4
Text Generation
•
1B
•
Updated
5 days ago
•
22
datasets
0
None public yet