Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Inference-native Tokenmaxxing Agent Harness for Loop Engineering
| Date | Stars |
|---|---|
| 2026-07-31 | 486 |
| 2026-08-04 | 489 |
| 2026-08-06 | 489 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<p align="center"> <img src="assets/inferoa-logo.svg" alt="Inferoa" width="420" /> </p> <p align="center"> <strong>Inference-native Tokenmaxxing Agent Harness for Loop Engineering</strong> </p> <p align="center"> <a href="https://github.com/agentic-in/inferoa">GitHub</a> · <a href="https://inferoa.agentic-in.ai/docs/intro">Docs</a> · <a href="https://inferoa.agentic-in.ai/blog/announcing-inferoa">Blog</a> </p> Prompting is no longer the whole interface. The frontier is **Loop Engineering**: give the model an objective, feedback, verification, memory, and tools, then let it self-correct until the work is proven. But every loop is also an inference workload. As turns accumulate, prompt prefixes drift, cache reuse collapses, stale evidence fills context, model routing gets harder, and serving choices start to matter. Inferoa is an **Inference-native Tokenmaxxing Agent Harness for Loop Engineering**: - **Inference-native**: the loop sees serving, routing, context windows, prefix cache, multimodal endpoints, and self-hosted model paths. - **Tokenmaxxing**: each turn is shaped to preserve cacheable prefixes, bound mutable context, expose token pressure, and pick the right inference path. - **Loop Engineering**: `/loop` runs durable recursive loops that inspect, edit, test, verify, decide, remember, and continue across loop tasks. ## Loop is All You Need <div align="center"> <p><strong>Loop Mode</strong></p> <img src="website/static/gif/loop.gif" alt="Inferoa loop mode" width="860" /> <p><strong>Intelligent Model Selection</strong></p> <img src="website/static/gif/model-selection.gif" alt="Inferoa intelligent model selection" width="860" /> <p><strong>Code Index</strong></p> <img src="website/static/gif/welcome.gif" alt="Inferoa code index" width="860" /> <p><strong>Plan Mode</strong></p> <img src="website/static/gif/plan.gif" alt="Inferoa plan mode" width="860" /> <p><strong>Loop Research</strong></p> <img src="website/static/gif/research.gif" alt="Inferoa research loop" width="860" /> </div> ## Why Inferoa Inferoa = **Infer**(Inference-native)**o**(Tokenmaxxing Loop)**a**(Agent Harness). <div align="center"> <img src="website/static/img/readme-why-inferoa.png" alt="Why Inferoa: Inference-native Loop Tokenmaxxing" width="860" /> </div> Inferoa gives that loop an inference-native runtime: - **Loop/rubric driven work**: `/loop` carries an objective across loop tasks, verification, decisions, recovery, and completion evidence instead of stopping after the next answer. - **Independent feedback surfaces**: plans, tests, tool results, research metrics, and completion evidence give the loop something concrete to improve against. - **Memory and context control**: compression, summaries, graph-shaped repo context, bounded history, and bounded tool output keep useful evidence in the window without letting stale state take over. - **Prefix-cache discipline**: prompt epochs, deterministic tool schemas, and bounded system sections protect reusable prefixes while the loop runs. - **Serving and routing remain visible**: model paths can respond to cost, safety, privacy, capability, session pressure, multimodal needs, and whether a self-hosted vLLM path is enough. ## The Tokenmaxxing Stack Inferoa is built on top of the vLLM ecosystem and extends tokenmaxxing across the inference stack: | Surface | Substrate | Inferoa role | Tokenmaxxing target | | --- | --- | --- | --- | | Loop Engineering | [Loop Mode](https://github.com/agentic-in/inferoa) | Recursive long-horizon loops, loop tasks, attempts, verification, decisions, completion evidence, and recovery | Keep the engineering loop running until the work is proven | | Agent Harness | [Inferoa](https://github.com/agentic-in/inferoa) | Sessions, tools, plans, loops, resources, evidence, and prefix-cache discipline | Give the loop a durable runtime while preserving reusable prompt prefixes | | Context Optimization | [CodeGraph](
Excerpt of 7,454 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:b097d643f3b6dd94, llm:Repository topics and description: 'agent, agent-harness, agentic-ai, harness-engineering, inference, loop-engineering, llm, tokenmaxxing' and description: 'Inference-native Tokenmaxxing Agent Harness for Loop Engineering'.
matched fp:b097d643f3b6dd94, llm:Repository topics and description: 'agent, agent-harness, agentic-ai, harness-engineering, inference, loop-engineering, llm, tokenmaxxing' and description: 'Inference-native Tokenmaxxing Agent Harness for Loop Engineering'.
matched fp:b097d643f3b6dd94, llm:Repository topics and description: 'agent, agent-harness, agentic-ai, harness-engineering, inference, loop-engineering, llm, tokenmaxxing' and description: 'Inference-native Tokenmaxxing Agent Harness for Loop Engineering'.