RedHatAI/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-FP8
Text Generation • 32B • Updated • 14 • 1
OpenSource and AI
SNLP: Layer-Parallel Inference via Structured Newton Corrections
S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation