feedd.AI
AIMarkTechPost · 10h ago

AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs

AMD released Instella-MoE-16B-A3B, a Mixture-of-Experts language model with 16 billion total parameters that activates only 2.8 billion per token. The model was trained on Instinct MI300X and MI325X GPUs with published weights from every training stage, data mixtures, configs, and inference code.

Read full story →
More from AI

The NVIDIA Transformer Engine uses fused GPU kernels and FP8 delayed scaling to optimize transformer workloads. The tutorial provides PyTorch code examples for training GPT-style causal language models with BF16 and FP8 precision formats.

01

An AI tool called Superapp generates native iOS apps written in Swift and SwiftUI from text prompts describing desired functionality. The system allows users to create applications without traditional coding knowledge or development experience.

02

Re-post-training upgraded DeepSeek-V4-Flash-0731 with improved agentic and coding capabilities while maintaining unchanged architecture and size. The model moved to public beta API on July 31, 2026, as the official release superseding the preview version.

03

Get feedd. daily

Top stories in your inbox every morning. Pick what you want.

No spam. Unsubscribe anytime.