feedd.AI
AIMarkTechPost · 28d ago

NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB

NVIDIA released three open embedding model checkpoints including an 8B version ranking first on RTEB with a 78.46 average NDCG@10 score. The 1B model was created through pruning and distillation from the 8B teacher, while the NVFP4 variant maintains 99% plus retrieval accuracy at up to 2x Blackwell throughput.

Read full story →
More from AI

Cerebras announced it is running OpenAI's GPT-5.6 Sol model at 750 output tokens per second on a new Ultrafast service tier, delivering up to 14 times faster processing than Standard. The Ultrafast tier launched as a limited preview in the OpenAI API for select customers on August 13, 2026.

01

Google released Gemini 3.7 Flash with a focus on coding and agentic workflows at half price through end of 2026, with input tokens costing $0.75 per million and output tokens $3.75 per million. The model arrives three weeks after Gemini 3.6 Flash, driven by developer feedback and algorithmic improvements aimed at reducing retries and manual oversight.

02

DeepSeek launched DeepSeek-V4-Pro, an updated flagship model for agentic workloads, and DeepSeek Harness v0.1, an open-source agent harness available on GitHub under MIT license. API pricing shifts to peak and off-peak rates beginning August 16 at 16:00 UTC, with substantially higher prices than current rates.

03

Get feedd. daily

Top stories in your inbox every morning. Pick what you want.

No spam. Unsubscribe anytime.