r/LocalLLaMA Project OntoPrune Ships Neuro-Symbolic Context Pruning for Local LLMs, 83% Fewer Tokens and 6.7x TTFT on CPU
Summary
Developer vigmarcarlo released OntoPrune, an MIT-licensed middleware that converts source code to RDF/SPARQL 'minimal context contracts' before feeding small LLMs. On a 300+ line Python file with Qwen 3B, it drops input tokens 83% (2,390 → 406), delivers 6.7x faster time-to-first-token on CPU (22.4s → 3.3s), and recorded zero invalid API calls versus one for the baseline. Ships as an MCP server, Python library and CLI.
Originally reported by github.com
Read the original article →Original headline: r/LocalLLaMA Project OntoPrune Ships Neuro-Symbolic Context Pruning for Local LLMs, 83% Fewer Tokens and 6.7x TTFT on CPU