Published event
DeveloperTools
Funding
1 source(s)
Are new OpenAI models getting better for coding?
Summary
Are new OpenAI models getting better for coding? This is the sixth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can’t control in the agent stack, how to measure whether your extensions are helping or hurting, and how to iterate toward better outcomes.
Why it matters
This Funding is relevant to the technology intelligence record because it involves OpenAI, GitHub. The source article should remain the factual reference for follow-up coverage.
Key facts
- This is the sixth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology.
- The series covers what you can and can’t control in the agent stack, how to measure whether your extensions are helping or hurting, and how to iterate toward better outcomes.
- A new model drops, the leaderboard says 92% on SWE-bench, and your timeline declares it “the best coding model.” You switch to it, run your agent on your codebase, and outcomes are… the same.
- The leaderboard said 92%, so what happened?
- In the previous article , we covered how to bootstrap agent knowledge for proprietary technology.
- That assumed you’d already picked a model and a harness.
Entities in this story
Technologies
Open Source→Related events