huggingface.co web signal

NVIDIA NeMo Retriever Paper Shows Agentic Retrieval Lifts nDCG@10 by 8.7 Points, Takes 160x Longer

NVIDIA RAG Agents ai-business

Summary

An NVIDIA NeMo Retriever paper reports agentic retrieval in a ReAct loop beats standard dense retrieval by 8.7 nDCG@10 points on average using the same embedding model, and holds the #1 spot on ViDoRe v3 and #2 on BRIGHT without reconfiguration. The pipeline consumes 764.1k input and 5.8k output tokens per query and averages 107.4 seconds per query vs 0.67 for standard retrieval. Opus 4.5 makes 9.2 search calls per query vs 2.4 for gpt-oss-120b, and the authors replace MCP with an in-process thread-safe retriever to cut overhead.