AI & ML interests
None defined yet.
Recent Activity
AtomGradient โ Bringing AI to the Edge
Research-driven on-device AI company building continual-learning and memory infrastructure.
Neural Imprint lets a model learn from local experience while base weights stay frozen and data never leaves the device.
๐ atomgradient.com
Research
Prism โ Cross-Domain Personal Data Integration on Consumer Hardware
Integrating finance, diet, mood, and reading data entirely on consumer Apple Silicon, producing emergent cross-domain insights with zero data leakage.
- ๐ 1.48x cross-domain insight emergence (IIR)
- ๐ 125.5x federation compression, zero data leakage
- โก 49.9 TPS real-time inference (35B on M2 Ultra)
ANE Batch Prefill โ On-Device Parallel LLM Inference
Fused matrix-vector kernels enabling concurrent ANE batch prefill + GPU decode on Apple Silicon for Qwen3.5 models.
- ๐ 11.3x ANE batch prefill speedup (268 tok/s)
- ๐ 79% power reduction for prefill component
- โฑ๏ธ <30 ms state transfer overhead
hybrid-ane-mlx-bench โ Disaggregated LLM Inference on Apple Silicon
Benchmarking CoreML ANE prefill + MLX GPU decode for Qwen3.5 on Apple Silicon, with four inference strategies compared.
- ๐ ANE prefill matches GPU at ~410 tokens
- ๐ 282x GPU power reduction during prefill
- ๐ 4 inference pipelines benchmarked
swift-qwen3-tts โ On-Device Text-to-Speech
Native Swift implementation of Qwen3 TTS 0.6B for real-time, on-device speech synthesis.
- ๐ฆ 67% model compression (2.35 GB โ 808 MB)
- ๐๏ธ Real-time synthesis (RTF 0.68x)
- ๐ 12 languages supported
Gemma-Prune โ On-Device Vision Language Model
Multi-stage compression pipeline for deploying Gemma 3 4B VLM on consumer hardware.
- ๐ฆ 25% model compression (2.8 GB โ 2.1 GB)
- ๐ 110 tok/s text generation
- ๐ผ๏ธ 3.4x image processing speedup
OptMLX โ MLX Memory Optimization Research
Exploring memory optimization techniques for the MLX framework on Apple Silicon.
- โก Up to 20x faster mmap loading
- ๐ Zero-copy model loading
- ๐ Comprehensive benchmarks
About
AtomGradient is an independent research group dedicated to making AI run efficiently on edge devices. Our research powers EchoStream AI โ a product line bringing on-device AI capabilities to real-world applications.
Edge AI ยท Privacy-First ยท Open Research