Skip to Main Content

Phoenix

Multi-agent GitHub issue resolution with baseline-aware safety checks

Image will load when scrolled into view
LLM AgentsSafety GatesEvaluationStatic Analysis

About the Project

Phoenix coordinates six specialized LLM agents to resolve GitHub issues from triage through pull-request creation, checking proposed changes against a baseline test run.

The evaluation includes a curated SWE-bench Lite slice and a pilot across real repositories. These bounded results should not be interpreted as full-benchmark performance or a guarantee of correctness on arbitrary codebases.

The preprint is available as 'Phoenix: Safe GitHub Issue Resolution via Multi-Agent LLMs', co-authored with Kipngeno Koech, Muhammad Adam, and Joao Barros.

Project Details

StatusPreprint
Role
Co-author & Architect
Stack
LLM Agents
Static Analysis Gates
Dynamic Validation
Python