facebookresearch/sweet_rl
quality grade D, 37 out of 100Benchmark and research code for the paper SWEET-RL Training Multi-Turn LLM Agents onCollaborative Reasoning Tasks
- stars
- 271
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Autonomous and multi-agent systems, agent frameworks, planning and tool use.
Signals: ai-agent, ai-agents, agents, autonomous-agents, agent-framework, multi-agent, multi-agent-systems, agentic
3,261 results
Benchmark and research code for the paper SWEET-RL Training Multi-Turn LLM Agents onCollaborative Reasoning Tasks
A new LLM solution for RTL code generation, achieving state-of-the-art performance in non-commercial solutions and outperforming GPT-3.5.
Chrome extension to filter your feed with LLM according to an explicitly stated and user-editable preference.
Semantic Kernel for Java. Integrate cutting-edge LLM technology quickly and easily into your Java based apps. See https://aka.ms/semantic-kernel.
Explorations into training LLMs to use clinical calculators from patient history, using open sourced models. Will start with Wells' Criteria
(ACL 2025 Main) Code for MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents https://www.arxiv.org/pdf/2503.01935
AI Agent Skills for Chinese Knowledge Workers: iMandalArt, FIRE, planning, and publishing workflows for Claude Code, Codex, and LLM agents.
Official Code Repository for the paper "Distilling LLM Agent into Small Models with Retrieval and Code Tools"
Embodied Agent Interface (EAI): Benchmarking LLMs for Embodied Decision Making (NeurIPS D&B 2024 Oral)
首个将99位加密KOL交易经验LLM蒸馏为可回测量化因子的开源项目 | First to distill 99 crypto KOL trading experience into backtestable quant factors via LLM
Official implementation for NeurIPS'24 paper: MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
DECeption with Evaluative Integrated Validation Engine (DECEIVE): Let an LLM do all the hard honeypot work!
LLMs sitting on a council together to decide, by consensus, who among them is the best.
Standalone agent runtime core with neutral LLM types, provider clients, tool loop, and extension protocols.
TapeAgents is a framework that facilitates all stages of the LLM Agent development lifecycle
Multi-Agent System Powered by LLMs for End-to-end Multimodal ML Automation
An automated Attack-with-Defense platform where LLM-powered agents compete in real-time.
Karpathy LLM based claude harness for PenetrationTesting / Bugbounty using obsidian
Extracting spatial and temporal world models from LLMs
A Test Project for a Network Security-oriented LLM Tool Emulating AutoGPT
Experimental library integrating LLM capabilities to support causal analyses
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
A minimal LLM-powered zero-day vulnerability scanner by AISLE.
LLM Agent and Evaluation Framework for Autonomous Penetration Testing
24,526 repositories in the index in total.