github.com via Reddit

r/LocalLLaMA Project OntoPrune Ships Neuro-Symbolic Context Pruning for Local LLMs, 83% Fewer Tokens and 6.7x TTFT on CPU

Open Source Inference ai-business

Summary

Developer vigmarcarlo released OntoPrune, an MIT-licensed middleware that converts source code to RDF/SPARQL 'minimal context contracts' before feeding small LLMs. On a 300+ line Python file with Qwen 3B, it drops input tokens 83% (2,390 → 406), delivers 6.7x faster time-to-first-token on CPU (22.4s → 3.3s), and recorded zero invalid API calls versus one for the baseline. Ships as an MCP server, Python library and CLI.