github.com via Reddit

OntoPrune Ships MIT Context Pruner for Local Coding LLMs

Open Source Inference ai-business

TL;DR

  • OntoPrune reports an 86% token cut on a Python benchmark (2,815 to 393) and 6.7x faster TTFT running Qwen 2.5 Coder 3B on Ollama.
  • The MIT-licensed middleware ships as an MCP server, Python API, and CLI, with tree-sitter parsing for Python, Dart, Java, and TypeScript.
  • Reported token reductions span 83 to 92.4% across the author's four-language benchmark set, with zero invalid methods on the Python run.

A new MIT-licensed middleware called OntoPrune claims it can cut the context a local coding LLM sees by 86% on a Python benchmark, from 2,815 tokens down to 393, and shave time-to-first-token on Qwen 2.5 Coder 3B over Ollama from 22.4 seconds to 3.3 seconds.

The README pitches the tool as a "Neuro-Symbolic Context Pruning Middleware for Local SLMs & Coding Agents (85% token reduction, 6.7x faster TTFT on CPU, 0% hallucinations)." It works, the author writes, by "isolating closed-world functional boundaries before attention computation," with reported reductions of 83 to 92.4% across four reference projects: Dart/Flutter at -91.8%, Java/Spring Boot at -92.4%, TypeScript/React at -92.4%.

The author, Vigmar Carlo, lists 45 passing tests and ships three entry points: an MCP server (ontoprune-mcp) targeting Claude Desktop and Cursor, a Python API, and a CLI with translate and check commands. Multi-language parsing goes through tree-sitter.

Every figure here is from the project's own README; the repo links no third-party replication. It lands into a visible wave of local-inference tooling the open-source tracker has been logging all quarter.