NVIDIA, Maryland: frozen medical AI learns from deployment
TL;DR
- A frozen LLM/VLM harness lifts medical-task accuracy by up to 34.2% over the base model with no retraining.
- The framework pairs three parts: a Skill mechanism, a Knowledge Memory of verified facts, and a Multimodal Knowledge Base of visual cases.
- Authors report gains across six benchmarks and four model variants, and claim transfer to non-medical domains.
Large medical AI models stop learning the day they ship. A new arXiv preprint from authors at NVIDIA and the University of Maryland proposes a workaround that leaves the model weights untouched.
The setup is model-agnostic. According to the abstract, the framework gives a frozen LLM or VLM three external faculties: "a Skill mechanism directing reasoning and tool application, a Knowledge Memory preserving verified facts from prior cases or authoritative sources, and a Multimodal Knowledge Base preserving visual examples while connecting retrieved cases to current images." Updates are gated by "an adaptive validation strategy" that only keeps a change if it improves new cases without harming prior ones.
The reported result: "performance gains reaching 34.2% above baseline on medical tasks," measured across six benchmarks spanning clinical diagnosis, workflow management, medical reasoning and visual reasoning, using four model variants. The authors also claim cross-model transferability and "applicability beyond medical domains."
The abstract does not name the specific base models tested, nor does it specify the hospitals, specialties, or data sources behind the deployment experience the system accumulates. It frames the motivation plainly: "clinical evidence and treatment protocols evolve," while fine-tuning "demands weight access and retraining" that most deployers do not have.
Originally reported by paper
Read the original article →Original headline: NVIDIA Paper: Frozen Medical AI Learns From Deployment, +34% on Clinical Tasks