English

Product LaunchesAmazon Web ServicesKiro

AWS Unveils System for Continuous Evaluation of AI Agent "Kiro" Prompts Using Real Conversation Data and LLMs

This article is a translation. Read the Japanese original

Amazon Web Services (AWS) has revealed a case study on its efforts to continuously evaluate the system prompts of its AI coding agent, "Kiro." By combining benchmark results with the analysis of thousands of conversation data points between internal developers and Kiro using Large Language Models (LLMs), the company is identifying opportunities for improvement in task completion and tool usage.

Because an AI agent's system prompt operates within every possible combination of foundation models, tools, codebases, tasks, and users, it is difficult to exhaustively verify all behaviors using a test suite. Furthermore, as improvements in LLM performance lead to more stringent interpretation of instructions, "prompt staleness"—where prompts tuned for older models inadvertently cause negative impacts—has become a challenge.

The evaluation framework introduced by AWS is centered around an "LLM Judge" that scores internal conversation sessions across 15 aspects of behavioral quality. The judge makes determinations based on specific evidence, such as explicit user dissatisfaction, conversation abandonment, or task interruption; it is designed to exclude ambiguous conversations from the evaluation to avoid excessive false positives.

The company uses this mechanism to compare changes to system prompts through A/B testing. Their operations include screening 27 potential changes to exclude those with adverse effects. Experimental results showed significant improvements: explicit dissatisfaction signals decreased by 5% in the Kiro CLI, and behavioral quality issues decreased by 20% in the Kiro IDE.

AWS also noted that when updating to a new model, excessive constraints intended for the old model can sometimes act as a bottleneck. In response, AWS treats model updates not merely as simple replacements, but as processes for re-tuning from scratch as a combination of the new model and updated prompts.

Sources

  1. 「システムプロンプト陳腐化」をどう防ぐか AWSの試行錯誤に見る「自社AIエージェント」運用のポイント (ITmedia AI+、2026-10-01)
  2. AWS公式ブログ