OntoPrune Ships MIT Context Pruner for Local Coding LLMs
TL;DR
- OntoPrune reports an 86% token cut on a Python benchmark (2,815 to 393) and 6.7x faster TTFT running Qwen 2.5 Coder 3B on Ollama.
- The MIT-licensed middleware ships as an MCP server, Python API, and CLI, with tree-sitter parsing for Python, Dart, Java, and TypeScript.
- Reported token reductions span 83 to 92.4% across the author's four-language benchmark set, with zero invalid methods on the Python run.
A new MIT-licensed middleware called OntoPrune claims it can cut the context a local coding LLM sees by 86% on a Python benchmark, from 2,815 tokens down to 393, and shave time-to-first-token on Qwen 2.5 Coder 3B over Ollama from 22.4 seconds to 3.3 seconds.
The README pitches the tool as a "Neuro-Symbolic Context Pruning Middleware for Local SLMs & Coding Agents (85% token reduction, 6.7x faster TTFT on CPU, 0% hallucinations)." It works, the author writes, by "isolating closed-world functional boundaries before attention computation," with reported reductions of 83 to 92.4% across four reference projects: Dart/Flutter at -91.8%, Java/Spring Boot at -92.4%, TypeScript/React at -92.4%.
The author, Vigmar Carlo, lists 45 passing tests and ships three entry points: an MCP server (ontoprune-mcp) targeting Claude Desktop and Cursor, a Python API, and a CLI with translate and check commands. Multi-language parsing goes through tree-sitter.
Every figure here is from the project's own README; the repo links no third-party replication. It lands into a visible wave of local-inference tooling the open-source tracker has been logging all quarter.
Originally reported by github.com
Read the original article →Original headline: r/LocalLLaMA Project OntoPrune Ships Neuro-Symbolic Context Pruning for Local LLMs, 83% Fewer Tokens and 6.7x TTFT on CPU