AI & ML interests

None defined yet.

Recent Activity

ShuhongWuย  updated a Space 1 day ago
AtomGradient/README
ShuhongWuย  updated a Space 6 months ago
AtomGradient/README
AlexWuKingย  updated a Space 6 months ago
AtomGradient/README
View all activity

Organization Card

AtomGradient โ€” Bringing AI to the Edge

Research-driven on-device AI company building continual-learning and memory infrastructure.

Neural Imprint lets a model learn from local experience while base weights stay frozen and data never leaves the device.

๐ŸŒ atomgradient.com


Research

Prism โ€” Cross-Domain Personal Data Integration on Consumer Hardware

Integrating finance, diet, mood, and reading data entirely on consumer Apple Silicon, producing emergent cross-domain insights with zero data leakage.

  • ๐Ÿ“ˆ 1.48x cross-domain insight emergence (IIR)
  • ๐Ÿ”’ 125.5x federation compression, zero data leakage
  • โšก 49.9 TPS real-time inference (35B on M2 Ultra)

[GitHub] ยท [Paper]


ANE Batch Prefill โ€” On-Device Parallel LLM Inference

Fused matrix-vector kernels enabling concurrent ANE batch prefill + GPU decode on Apple Silicon for Qwen3.5 models.

  • ๐Ÿš€ 11.3x ANE batch prefill speedup (268 tok/s)
  • ๐Ÿ”‹ 79% power reduction for prefill component
  • โฑ๏ธ <30 ms state transfer overhead

[GitHub] ยท [Paper]


hybrid-ane-mlx-bench โ€” Disaggregated LLM Inference on Apple Silicon

Benchmarking CoreML ANE prefill + MLX GPU decode for Qwen3.5 on Apple Silicon, with four inference strategies compared.

  • ๐Ÿ”„ ANE prefill matches GPU at ~410 tokens
  • ๐Ÿ”‹ 282x GPU power reduction during prefill
  • ๐Ÿ“Š 4 inference pipelines benchmarked

[GitHub] ยท [Paper]


swift-qwen3-tts โ€” On-Device Text-to-Speech

Native Swift implementation of Qwen3 TTS 0.6B for real-time, on-device speech synthesis.

  • ๐Ÿ“ฆ 67% model compression (2.35 GB โ†’ 808 MB)
  • ๐ŸŽ™๏ธ Real-time synthesis (RTF 0.68x)
  • ๐ŸŒ 12 languages supported

[GitHub] ยท [Paper]


Gemma-Prune โ€” On-Device Vision Language Model

Multi-stage compression pipeline for deploying Gemma 3 4B VLM on consumer hardware.

  • ๐Ÿ“ฆ 25% model compression (2.8 GB โ†’ 2.1 GB)
  • ๐Ÿ“ 110 tok/s text generation
  • ๐Ÿ–ผ๏ธ 3.4x image processing speedup

[GitHub] ยท [Paper]


OptMLX โ€” MLX Memory Optimization Research

Exploring memory optimization techniques for the MLX framework on Apple Silicon.

  • โšก Up to 20x faster mmap loading
  • ๐Ÿ”„ Zero-copy model loading
  • ๐Ÿ“Š Comprehensive benchmarks

[GitHub] ยท [Paper]


About

AtomGradient is an independent research group dedicated to making AI run efficiently on edge devices. Our research powers EchoStream AI โ€” a product line bringing on-device AI capabilities to real-world applications.

Edge AI ยท Privacy-First ยท Open Research

datasets 0

None public yet