Can AI agents build real Stripe integrations? We built a benchmark to find out (opens in new tab)
State-of-the-art LLM agents can complete many scoped coding tasks, but fully autonomous software engineering remains difficult because real projects require long-term planning, persistent state, debugging, and end-to-end validation. Stripe evaluated this gap through a benchmark of realistic backend, frontend, database, and browser-based integration tasks. The results were stronger than expected: agents demonstrated substantial full-stack capability, but still struggled with ambiguity and the judgment required to distinguish genuine failures from bad test inputs.
Building the Stripe Integration Benchmark
- Stripe created 11 environments based on real integration challenges, including Checkout migrations and Billing API modeling.
- Each environment included:
- A complete codebase, database, scripts, and test Stripe credentials.
- Deterministic graders using API calls, automated browser tests, or inspection of Stripe objects.
- A consistent agent harness with terminal, browser, and Stripe-specific search tools through MCP.
- Challenges were divided into:
- Backend-only tasks: SDK upgrades, API changes, and database migrations.
- Full-stack tasks: Coordinated server and client changes requiring browser verification.
- Gym problem sets: Focused exercises testing deep knowledge of features such as Checkout and subscriptions.
Stronger-than-Expected Agent Performance
- The benchmark intentionally used fewer, harder tasks designed to expose weaknesses.
- Agents successfully:
- Navigated browser interfaces.
- Debugged live issues.
- Worked with underdocumented API behavior.
- Continued productively across long interactions, with top runs averaging 63 turns.
- Claude Opus 4.5 achieved a 92% average score across four full-stack tasks.
- GPT-5.2 achieved a 73% average score across two gym problem sets.
- In a migration from Card Element to Checkout, an agent completed and verified a test purchase using Link, despite no payment method being specified.
Reverse-Engineering Checkout Configurations
- A Checkout gym task required agents to infer API parameters from 20 prebuilt Checkout UIs.
- Agents had to:
- Inspect products and quantities shown in each session.
- Locate matching product IDs through the Products API.
- Identify shipping costs, custom fields, tax settings, and other customizations.
- Translate those details into valid Checkout Session parameters.
- Agents provided more than 80% of the correct parameters.
- The best-performing agent recognized that one UI’s color options were hidden behind an interactive dropdown, explored the control, and included the missing values.
Remaining Challenges with Ambiguity
- Agents struggled when evaluation situations required judgment rather than straightforward implementation.
- In SDK upgrade tasks, some agents supplied nonexistent Stripe data, received expected 400 errors, and treated those responses as evidence that their implementation was broken.
- This illustrates a broader limitation: successful autonomous engineering requires not only writing code, but also designing meaningful tests, interpreting failures correctly, and validating behavior against realistic system state.
The benchmark suggests that agents are increasingly capable of substantial Stripe integration work, including full-stack implementation and browser-based verification. However, reliable autonomy will require better handling of ambiguity, realistic test data, persistent project state, and rigorous end-to-end validation.