Published event
ArtificialIntelligence Research 1 source(s)

Your Agent Aced the Task. Will It Do It Again?

Updated September 26, 2026 · 2:44 PM · source date September 15, 2026

Summary

Your Agent Aced the Task. Will It Do It Again? Enterprise Article Published September 15, 2026 Upvote 112 Evelyn Duesterwald evduester ibm-research Lilian Ngweta lilianngweta ibm-research Vatche Isahagian Vatche ibm-research Jayaram Radhakrishnan jayaramkr ibm-research Vinod Muthusamy vinodmut ibm-research Gaodan Fang gaodan-fang ibm-research Ashwath Vaithinathan Aravindan ashwath-vaithina ibm-research Punleuk Oum illeatmyhat ibm-research G Thomas gsthomasx ibm-research Merve Unuvar mrvnvr ibm-research Ayhan Sebin ayhansebin ibm-research Michał Ulewicz Michal ibm-research Your agent works in rehearsal, but during the live demo, it takes a different path and fails the same task. In production, it is a reliability problem: a workflow that succeeded once may fail the next time a user makes the same request.

Why it matters

This Research is relevant to the technology intelligence record because it involves GitHub, Intel, gpt-4.1, gpt-oss. The source article should remain the factual reference for follow-up coverage.

Key facts
  • Enterprise Article Published September 15, 2026 Upvote 112 Evelyn Duesterwald evduester ibm-research Lilian Ngweta lilianngweta ibm-research Vatche Isahagian Vatche ibm-research Jayaram Radhakrishnan jayaramkr ibm-research Vinod Muthusamy vinodmut ibm-research Gaodan Fang gaodan-fang ibm-research Ashwath Vaithinathan Aravindan ashwath-vaithina ibm-research Punleuk Oum illeatmyhat ibm-research G Thomas gsthomasx ibm-research Merve Unuvar mrvnvr ibm-research Ayhan Sebin ayhansebin ibm-research Michał Ulewicz Michal ibm-research Your agent works in rehearsal, but during the live demo, it takes a different path and fails the same task.
  • In production, it is a reliability problem: a workflow that succeeded once may fail the next time a user makes the same request.
  • For mission-critical work, such as reconciling a financial transaction or checking a contract for an obligation, that can be a showstopper.
  • Most benchmarks hide this variability behind an average.
  • On AppWorld, a ReAct agent using GPT-4.1 succeeded on 77.4% of runs across five repetitions.
  • But it succeeded in all five runs for only 53.0% of tasks — a 24.4-point consistency gap .
Entities in this story
Related events