[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"kb-article-ai-agent-evaluation-best-practices-from-google-experts-en":3,"ArticleBody_9GuCgpIm3yscbNavCVjN2ODs8bz5hCfFlsPtXBI94":225},{"article":4,"relatedArticles":195,"locale":66},{"id":5,"title":6,"slug":7,"content":8,"htmlContent":9,"excerpt":10,"category":11,"tags":12,"metaDescription":10,"wordCount":13,"readingTime":14,"publishedAt":15,"sources":16,"sourceCoverage":58,"transparency":60,"seo":63,"language":66,"featuredImage":67,"featuredImageCredit":68,"isFreeGeneration":72,"trendSlug":7,"trendSnapshot":73,"niche":81,"geoTakeaways":84,"geoFaq":93,"entities":103},"6a64468f07f0903672a87e6f","AI Agent Evaluation Best Practices from Google Experts","ai-agent-evaluation-best-practices-from-google-experts","Modern AI is moving from single-shot chat to agents that plan, call tools, and run workflows across critical systems.[1][3] Evaluating them like static QA models misses whether they used the right tools, respected policies, or followed safe sequences.[1][2] As agents enter finance, supply chains, and HR, [evaluation becomes a governance obligation](\u002Farticle\u002Freliability-focused-evaluation-methods-for-agentic-ai-systems).[3][7]  \n\n**Key takeaway:** If agents can act, evaluation must inspect how they reason and operate, not just what they say.[1][3]  \n\n---\n\n## 1. Why Modern AI Agents Need a New Evaluation Mindset\n\n[AI agents](\u002Fentities\u002F695e94e819d266277e14e030-ai-agents) now orchestrate multi-step workflows, call APIs, and update systems autonomously.[1][3] A binary right\u002Fwrong label no longer shows where a chain of decisions failed.[1][2]  \n\nKey risks:[1][3][6][7]  \n\n- **Silent failure:**  \n  - Agent returns a correct-looking answer via wrong, unsafe, or non-compliant steps.  \n- **Large blast radius:**  \n  - [SAP](\u002Fentities\u002F69600eff19d266277e14fa15-sap)’s [Joule](\u002Fentities\u002F6a5fb57e457d25046952b83e-joule) Agents report up to 75% less time on complex workflows and aim for 30%+ higher efficiency across systems.[6][8]  \n  - When such agents misstep in sourcing, finance, or supply chains, a single flawed intermediate action can create costly, opaque side effects.[6][7]  \n- **Vibe-based testing:**  \n  - Many teams rely on ad hoc manual runs, eyeballing transcripts, and sporadic tool use.[4]  \n  - One Reddit practitioner described having no formal framework even after exploring eval platforms.[4][3]  \n\n**Key point:** Without shifting from “answer quality” to “process and interaction quality,” you will ship agents that demo well yet fail silently in production.[1][3]  \n\n---\n\n## 2. A [Google](\u002Fentities\u002F6960e32919d266277e1504f6-google)-Inspired Framework for Evaluating AI Agents End-to-End\n\nGoogle distinguishes between evaluating the [final response](https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FResponse_Boat_%E2%80%93_Medium) and the [trajectory](https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FTrajectory).[2]  \n\n- **Final response:** Did the agent achieve the user’s goal correctly and helpfully?[1][2]  \n- **Trajectory:** Was the sequence of reasoning and tool calls appropriate, efficient, and safe?[1][2]  \n\nWith Google’s Gen AI evaluation service, a single Vertex AI SDK call can run an agent and return both types of metrics for Agent Engine templates, [LangChain](\u002Fentities\u002F6960e29a19d266277e150441-langchain)-based agents, or custom functions.[2]  \n\n**Key takeaway:** Treat “final vs trajectory” as two contracts your agent must satisfy, both enforced by evaluation.[1][2]  \n\n### 2.1 Build a Structured Quality Gate\n\nA practical quality gate includes:[1][2]  \n\n1. **Define success precisely**  \n   - E.g., “Book a multi-leg flight that satisfies all constraints with zero booking errors,” not “be helpful.”[1]  \n\n2. **Instrument the agent**  \n   - Log trajectory steps, tools and parameters, tool outputs, and full conversations.[1]  \n\n3. **Score final outputs**  \n   - Correctness, helpfulness, goal completion via human raters or LLM-as-judge.[1][2]  \n\n4. **Score each trajectory step**  \n   - **Validity:** right tool, parameters, no policy violations.  \n   - **Efficiency:** no unnecessary calls, loops, redundant retrievals.  \n   - **Safety:** prompt-injection resistance, guardrail adherence.[1][2]  \n\n**Key point:** You cannot diagnose failures—or prove compliance—without full interaction-level logging across prompts, thoughts, and tools.[1]  \n\nA useful way to visualize this is as an end-to-end workflow, from defining success through gating deployments, with evaluation signals feeding into your governance processes.\n\n```mermaid\nflowchart TB\n    title End-to-End AI Agent Evaluation Workflow\n    A[Define success] --> B[Instrument agent]\n    B --> C[Run eval sets]\n    C --> D[Score responses]\n    C --> E[Score trajectories]\n    D --> F[Detect anomalies]\n    E --> F\n    F --> G[Gate deployments]\n    classDef success fill:#22c55e,stroke:#14532d,color:#ffffff\n    classDef warning fill:#f59e0b,stroke:#92400e,color:#000000\n    classDef danger fill:#ef4444,stroke:#7f1d1d,color:#ffffff\n    classDef info fill:#3b82f6,stroke:#1e3a8a,color:#ffffff\n    class A success\n    class B info\n    class C info\n    class D success\n    class E warning\n    class F warning\n    class G danger\n```\n\n### 2.2 Tie Evaluation to Business Value\n\nReal ROI comes from automating workflows in domains like healthcare and finance, not better chit-chat.[3][6] Metrics should reflect:[3][7]  \n\n- Cycle-time reduction for key workflows  \n- Error-rate reduction vs. legacy processes  \n- Workflow completion and handoff success  \n\nSAP’s Joule platform positions agents as drivers of enterprise productivity and cost savings via cross-system workflows under strong governance.[6][7]  \n\n**Business lens:** If dashboards show only language scores and not “average days-to-close reduced,” you are misaligned with executives.[3][7]  \n\n### 2.3 A Tiered Maturity Model\n\nInspired by Google and other platforms, grow evaluation in tiers:[2][7][8]  \n\n1. **Tier 0 – Manual checks**  \n   - Small test suites, human review of final responses only.  \n\n2. **Tier 1 – Automated dual evaluation**  \n   - Use Gen AI eval or similar to score final and trajectory metrics on every experiment.[2]  \n\n3. **[Tier 2](\u002Fentities\u002F69ebce4ae1ca17caac376494-tier-2) – Enterprise-grade gates**  \n   - Link evaluation to governance workflows, [security approvals](\u002Farticle\u002Fmicrosoft-agent-365-security-features-for-enterprise-grade-agentic-ai), and compliance before agents touch live finance, HR, or supply-chain systems.[7][8]  \n\n**Maturity signal:** Promotion to production should require passing final-response and trajectory thresholds at Tier 2—not just “looks good.”[1][2]  \n\n---\n\n## 3. Operationalizing Agent Evaluation in Real Teams\n\nEngineering leaders should integrate agent evaluation into observability and incident pipelines. Google recommends treating evaluation events as first-class telemetry alongside latency and error logs, especially for tool-orchestrating agents on platforms like Agent Engine.[1][2][7]  \n\nExample:[6][7]  \n\n- A procurement agent at a 30-person manufacturer began over-ordering after a supplier API change.  \n- Only detailed tool-call logs plus workflow metrics exposed the regression before it hit the P&L.  \n\n**Key takeaway:** Route evaluation scores, trajectory anomalies, and safety violations into the same alerting stack used for microservices health.[1][7]  \n\n### 3.1 Curated Eval Sets from Real Traffic\n\nBuild eval datasets from real interactions, especially in autonomy-heavy domains like sourcing and finance:[3][6]  \n\n- Mine logs for tricky, multi-system workflows  \n- Capture prompt-injection attempts and escalation scenarios  \n- Refresh eval sets as behavior and tools change  \n\n### 3.2 Shared Ownership and Continuous Recalibration\n\nEvaluation cannot sit only with ML engineers.[3][7][8]  \n\n- Product and domain experts define success, policy constraints, and acceptable risk.  \n- As infrastructure and models evolve—similar to [OpenAI](\u002Fentities\u002F695e3c6f19d266277e14dd48-openai)’s full-stack optimization for [Jalapeño](\u002Fentities\u002F6a3dc9fac460e8b42cddc21c-jalapeno), where hardware, serving stack, and models co-evolve—revisit metrics, thresholds, and trajectory templates.[9][10]  \n- New tools or latency envelopes change what “good” looks like.  \n\n---\n\n## Conclusion\n\n**Key point:** Treat your evaluation suite as a living artifact, updated as your models, tools, and business requirements evolve—and make both final answers and trajectories pass explicit, business-aligned gates before agents act in critical systems.","\u003Cp>Modern AI is moving from single-shot chat to agents that plan, call tools, and run workflows across critical systems.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa> Evaluating them like static QA models misses whether they used the right tools, respected policies, or followed safe sequences.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa> As agents enter finance, supply chains, and HR, \u003Ca href=\"\u002Farticle\u002Freliability-focused-evaluation-methods-for-agentic-ai-systems\" class=\"internal-link\">evaluation becomes a governance obligation\u003C\u002Fa>.\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Key takeaway:\u003C\u002Fstrong> If agents can act, evaluation must inspect how they reason and operate, not just what they say.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003C\u002Fp>\n\u003Chr>\n\u003Ch2>1. Why Modern AI Agents Need a New Evaluation Mindset\u003C\u002Fh2>\n\u003Cp>\u003Ca href=\"\u002Fentities\u002F695e94e819d266277e14e030-ai-agents\">AI agents\u003C\u002Fa> now orchestrate multi-step workflows, call APIs, and update systems autonomously.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa> A binary right\u002Fwrong label no longer shows where a chain of decisions failed.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>Key risks:\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Silent failure:\u003C\u002Fstrong>\n\u003Cul>\n\u003Cli>Agent returns a correct-looking answer via wrong, unsafe, or non-compliant steps.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Large blast radius:\u003C\u002Fstrong>\n\u003Cul>\n\u003Cli>\u003Ca href=\"\u002Fentities\u002F69600eff19d266277e14fa15-sap\">SAP\u003C\u002Fa>’s \u003Ca href=\"\u002Fentities\u002F6a5fb57e457d25046952b83e-joule\">Joule\u003C\u002Fa> Agents report up to 75% less time on complex workflows and aim for 30%+ higher efficiency across systems.\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>When such agents misstep in sourcing, finance, or supply chains, a single flawed intermediate action can create costly, opaque side effects.\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Vibe-based testing:\u003C\u002Fstrong>\n\u003Cul>\n\u003Cli>Many teams rely on ad hoc manual runs, eyeballing transcripts, and sporadic tool use.\u003Ca href=\"#source-4\" class=\"citation-link\" title=\"View source [4]\">[4]\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>One Reddit practitioner described having no formal framework even after exploring eval platforms.\u003Ca href=\"#source-4\" class=\"citation-link\" title=\"View source [4]\">[4]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cstrong>Key point:\u003C\u002Fstrong> Without shifting from “answer quality” to “process and interaction quality,” you will ship agents that demo well yet fail silently in production.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003C\u002Fp>\n\u003Chr>\n\u003Ch2>2. A \u003Ca href=\"\u002Fentities\u002F6960e32919d266277e1504f6-google\">Google\u003C\u002Fa>-Inspired Framework for Evaluating AI Agents End-to-End\u003C\u002Fh2>\n\u003Cp>Google distinguishes between evaluating the \u003Ca href=\"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FResponse_Boat_%E2%80%93_Medium\" class=\"wiki-link\" target=\"_blank\" rel=\"noopener\">final response\u003C\u002Fa> and the \u003Ca href=\"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FTrajectory\" class=\"wiki-link\" target=\"_blank\" rel=\"noopener\">trajectory\u003C\u002Fa>.\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Final response:\u003C\u002Fstrong> Did the agent achieve the user’s goal correctly and helpfully?\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Trajectory:\u003C\u002Fstrong> Was the sequence of reasoning and tool calls appropriate, efficient, and safe?\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>With Google’s Gen AI evaluation service, a single Vertex AI SDK call can run an agent and return both types of metrics for Agent Engine templates, \u003Ca href=\"\u002Fentities\u002F6960e29a19d266277e150441-langchain\">LangChain\u003C\u002Fa>-based agents, or custom functions.\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Key takeaway:\u003C\u002Fstrong> Treat “final vs trajectory” as two contracts your agent must satisfy, both enforced by evaluation.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fp>\n\u003Ch3>2.1 Build a Structured Quality Gate\u003C\u002Fh3>\n\u003Cp>A practical quality gate includes:\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fp>\n\u003Col>\n\u003Cli>\n\u003Cp>\u003Cstrong>Define success precisely\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>E.g., “Book a multi-leg flight that satisfies all constraints with zero booking errors,” not “be helpful.”\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\n\u003Cp>\u003Cstrong>Instrument the agent\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Log trajectory steps, tools and parameters, tool outputs, and full conversations.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\n\u003Cp>\u003Cstrong>Score final outputs\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Correctness, helpfulness, goal completion via human raters or LLM-as-judge.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\n\u003Cp>\u003Cstrong>Score each trajectory step\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Validity:\u003C\u002Fstrong> right tool, parameters, no policy violations.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Efficiency:\u003C\u002Fstrong> no unnecessary calls, loops, redundant retrievals.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Safety:\u003C\u002Fstrong> prompt-injection resistance, guardrail adherence.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>\u003Cstrong>Key point:\u003C\u002Fstrong> You cannot diagnose failures—or prove compliance—without full interaction-level logging across prompts, thoughts, and tools.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>A useful way to visualize this is as an end-to-end workflow, from defining success through gating deployments, with evaluation signals feeding into your governance processes.\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-mermaid\">flowchart TB\n    title End-to-End AI Agent Evaluation Workflow\n    A[Define success] --&gt; B[Instrument agent]\n    B --&gt; C[Run eval sets]\n    C --&gt; D[Score responses]\n    C --&gt; E[Score trajectories]\n    D --&gt; F[Detect anomalies]\n    E --&gt; F\n    F --&gt; G[Gate deployments]\n    classDef success fill:#22c55e,stroke:#14532d,color:#ffffff\n    classDef warning fill:#f59e0b,stroke:#92400e,color:#000000\n    classDef danger fill:#ef4444,stroke:#7f1d1d,color:#ffffff\n    classDef info fill:#3b82f6,stroke:#1e3a8a,color:#ffffff\n    class A success\n    class B info\n    class C info\n    class D success\n    class E warning\n    class F warning\n    class G danger\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch3>2.2 Tie Evaluation to Business Value\u003C\u002Fh3>\n\u003Cp>Real ROI comes from automating workflows in domains like healthcare and finance, not better chit-chat.\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa> Metrics should reflect:\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Cycle-time reduction for key workflows\u003C\u002Fli>\n\u003Cli>Error-rate reduction vs. legacy processes\u003C\u002Fli>\n\u003Cli>Workflow completion and handoff success\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>SAP’s Joule platform positions agents as drivers of enterprise productivity and cost savings via cross-system workflows under strong governance.\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Business lens:\u003C\u002Fstrong> If dashboards show only language scores and not “average days-to-close reduced,” you are misaligned with executives.\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fp>\n\u003Ch3>2.3 A Tiered Maturity Model\u003C\u002Fh3>\n\u003Cp>Inspired by Google and other platforms, grow evaluation in tiers:\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003C\u002Fp>\n\u003Col>\n\u003Cli>\n\u003Cp>\u003Cstrong>Tier 0 – Manual checks\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Small test suites, human review of final responses only.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\n\u003Cp>\u003Cstrong>Tier 1 – Automated dual evaluation\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Use Gen AI eval or similar to score final and trajectory metrics on every experiment.\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\n\u003Cp>\u003Cstrong>\u003Ca href=\"\u002Fentities\u002F69ebce4ae1ca17caac376494-tier-2\">Tier 2\u003C\u002Fa> – Enterprise-grade gates\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Link evaluation to governance workflows, \u003Ca href=\"\u002Farticle\u002Fmicrosoft-agent-365-security-features-for-enterprise-grade-agentic-ai\" class=\"internal-link\">security approvals\u003C\u002Fa>, and compliance before agents touch live finance, HR, or supply-chain systems.\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>\u003Cstrong>Maturity signal:\u003C\u002Fstrong> Promotion to production should require passing final-response and trajectory thresholds at Tier 2—not just “looks good.”\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fp>\n\u003Chr>\n\u003Ch2>3. Operationalizing Agent Evaluation in Real Teams\u003C\u002Fh2>\n\u003Cp>Engineering leaders should integrate agent evaluation into observability and incident pipelines. Google recommends treating evaluation events as first-class telemetry alongside latency and error logs, especially for tool-orchestrating agents on platforms like Agent Engine.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>Example:\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>A procurement agent at a 30-person manufacturer began over-ordering after a supplier API change.\u003C\u002Fli>\n\u003Cli>Only detailed tool-call logs plus workflow metrics exposed the regression before it hit the P&amp;L.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cstrong>Key takeaway:\u003C\u002Fstrong> Route evaluation scores, trajectory anomalies, and safety violations into the same alerting stack used for microservices health.\u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003C\u002Fp>\n\u003Ch3>3.1 Curated Eval Sets from Real Traffic\u003C\u002Fh3>\n\u003Cp>Build eval datasets from real interactions, especially in autonomy-heavy domains like sourcing and finance:\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Mine logs for tricky, multi-system workflows\u003C\u002Fli>\n\u003Cli>Capture prompt-injection attempts and escalation scenarios\u003C\u002Fli>\n\u003Cli>Refresh eval sets as behavior and tools change\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>3.2 Shared Ownership and Continuous Recalibration\u003C\u002Fh3>\n\u003Cp>Evaluation cannot sit only with ML engineers.\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Product and domain experts define success, policy constraints, and acceptable risk.\u003C\u002Fli>\n\u003Cli>As infrastructure and models evolve—similar to \u003Ca href=\"\u002Fentities\u002F695e3c6f19d266277e14dd48-openai\">OpenAI\u003C\u002Fa>’s full-stack optimization for \u003Ca href=\"\u002Fentities\u002F6a3dc9fac460e8b42cddc21c-jalapeno\">Jalapeño\u003C\u002Fa>, where hardware, serving stack, and models co-evolve—revisit metrics, thresholds, and trajectory templates.\u003Ca href=\"#source-9\" class=\"citation-link\" title=\"View source [9]\">[9]\u003C\u002Fa>\u003Ca href=\"#source-10\" class=\"citation-link\" title=\"View source [10]\">[10]\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>New tools or latency envelopes change what “good” looks like.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Chr>\n\u003Ch2>Conclusion\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>Key point:\u003C\u002Fstrong> Treat your evaluation suite as a living artifact, updated as your models, tools, and business requirements evolve—and make both final answers and trajectories pass explicit, business-aligned gates before agents act in critical systems.\u003C\u002Fp>\n","Modern AI is moving from single-shot chat to agents that plan, call tools, and run workflows across critical systems.[1][3] Evaluating them like static QA models misses whether they used the right too...","trend-radar",[],941,5,"2026-07-25T05:22:41.381Z",[17,22,26,30,34,38,42,46,50,54],{"title":18,"url":19,"summary":20,"type":21},"A methodical approach to agent evaluation: Building a robust quality gate","https:\u002F\u002Fcloud.google.com\u002Fblog\u002Ftopics\u002Fdevelopers-practitioners\u002Fa-methodical-approach-to-agent-evaluation","A methodical approach to agent evaluation: Building a robust quality gate\n\nNovember 17, 2025\n\n##### Hugo Selbie\n\nStaff Customer & Partner Solutions Engineer, Google\n\n##### Try Gemini Enterprise Busine...","kb",{"title":23,"url":24,"summary":25,"type":21},"Evaluate Gen AI agents","https:\u002F\u002Fdocs.cloud.google.com\u002Fgemini-enterprise-agent-platform\u002Fmodels\u002Fevaluation-agents","## Preview\n\nThis feature is subject to the \"Pre-GA Offerings Terms\" in the General Service Terms section of the [Service Specific Terms](https:\u002F\u002Fdocs.cloud.google.com\u002Fterms\u002Fservice-terms#1). Pre-GA fe...",{"title":27,"url":28,"summary":29,"type":21},"Best Practices to Navigate the Complexities of Evaluating AI Agents","https:\u002F\u002Fgalileo.ai\u002Fblog\u002Fevaluating-ai-agents-best-practices","Apr 17, 2025 — Conor Bronsdon\n\nBest Practices to Navigate the Complexities of Evaluating AI Agents\n\nAI is moving from simple conversation tools to robust systems driving automation in various industri...",{"title":31,"url":32,"summary":33,"type":21},"For people out there making AI agents, how are you evaluating the performance of your agent?","https:\u002F\u002Fwww.reddit.com\u002Fr\u002FAI_Agents\u002Fcomments\u002F1k1mmb1\u002Ffor_people_out_there_making_ai_agents_how_are_you\u002F","Hey everyone - I've recently realized testing AI agents beyond manual QA is not trivial, and I don't have a framework for properly testing my agent. Looked at LangSmith and Arize, and it seems like th...",{"title":35,"url":36,"summary":37,"type":21},"SAP presents Joule AI Agents: An extensive portfolio of AI agents","https:\u002F\u002Fwww.techzine.eu\u002Fnews\u002Fanalytics\u002F131606\u002Fsap-presents-new-joule-ai-agents-within-the-business-suite\u002F","SAP presents an extensive portfolio of Joule Agents that automatically collaborate with people. The AI agents provide smart insights and help users, enabling organizations to operate more efficiently....",{"title":39,"url":40,"summary":41,"type":21},"Joule Agents and Joule Assistants","https:\u002F\u002Fwww.sap.com\u002Fproducts\u002Fartificial-intelligence\u002Fai-agents.html","Accelerate outcomes with context-aware agents and assistants that know your work and your business. \n\nJoule Agents are AI agents with business process expertise that automate workflows at scale. Joule...",{"title":43,"url":44,"summary":45,"type":21},"Capture business-wide AI value with speed and confidence","https:\u002F\u002Fwww.sap.com\u002Fproducts\u002Fartificial-intelligence.html","Capture business-wide AI value with speed and confidence\n\nJoule brings assistants and agents together in a unified workspace that turns intent into autonomous action—running end-to-end workflows acros...",{"title":47,"url":48,"summary":49,"type":21},"What is Joule?","https:\u002F\u002Fpages.community.sap.com\u002Ftopics\u002Fjoule","Joule is an AI solution that helps teams act with clarity and collaborate without silos. With AI agents for all core functions, powered by SAP business process expertise, your AI strategy scales faste...",{"title":51,"url":52,"summary":53,"type":21},"OpenAI's Jalapeño: AI Designed Inference Chip for LLMs","https:\u002F\u002Fwww.linkedin.com\u002Fposts\u002Frichard-ho-chips_openai-and-broadcom-unveil-llm-optimized-activity-7475540055822901248-_988","Richard Ho\n1mo\n\nWhen we started Jalapeño, the question was not “how do we build another AI accelerator?” It was: what should an inference chip look like if it is designed around the way modern LLMs ac...",{"title":55,"url":56,"summary":57,"type":21},"OpenAI Jalapeño Chip Explained: What OpenAI's First Custom Inference ASIC Means for GPU Cloud (2026)","https:\u002F\u002Fwww.spheron.network\u002Fblog\u002Fopenai-jalapeno-chip-gpu-cloud-inference-2026\u002F","OpenAI's Jalapeño chip is a custom LLM inference ASIC built with Broadcom, targeting a 10 GW infrastructure commitment through 2029. It is real, it is significant at OpenAI's scale, and it has no bear...",{"totalSources":59},10,{"generationDuration":61,"kbQueriesCount":59,"confidenceScore":62,"sourcesCount":59},128921,100,{"metaTitle":64,"metaDescription":65},"AI Agent Evaluation: Governance, Safety & Best Practices","Move beyond accuracy metrics—learn to evaluate AI agents' reasoning, tool use, and workflows for safety and compliance. Get practical frameworks to cut deployme","en","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1586282023639-dbd68e65a9fe?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxhZ2VudCUyMGV2YWx1YXRpb24lMjBiZXN0JTIwcHJhY3RpY2VzfGVufDF8MHx8fDE3ODQ5NTY1NTl8MA&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60",{"photographerName":69,"photographerUrl":70,"unsplashUrl":71},"Markus Winkler","https:\u002F\u002Funsplash.com\u002F@markuswinkler?utm_source=coreprose&utm_medium=referral","https:\u002F\u002Funsplash.com\u002Fphotos\u002Fwhite-printer-paper-on-macbook-pro-yp6tMFz2qy8?utm_source=coreprose&utm_medium=referral",true,{"score":62,"type":74,"sourceCount":75,"topSourceDomains":76,"detectedAt":80,"mentionsLast7Days":75},"spiking",7,[77,78,79],"startuphub.ai","solutionsreview.com","patmcguinness.substack.com","2026-07-25T03:04:38.382Z",{"key":82,"name":83,"nameEn":83},"ai-engineering","AI Engineering & LLM Ops",[85,87,89,91],{"text":86},"Modern agents must be evaluated on both final responses and trajectories; Google-style evaluation returns both types of metrics in a single Vertex AI SDK call for Agent Engine, LangChain, or custom agents.",{"text":88},"Instrumentation and logging of every trajectory step (tool called, parameters, outputs, and intermediate reasoning) are mandatory; without full interaction-level logs you cannot prove compliance or diagnose silent failures.",{"text":90},"Enterprise deployments require tiered gates: Tier 0 (manual checks), Tier 1 (automated dual evaluation), Tier 2 (enterprise-grade gates). Promotion to production must require passing final-response and trajectory thresholds at Tier 2.",{"text":92},"Agents can drastically change business outcomes: SAP’s Joule reports up to 75% less time on complex workflows and targets 30%+ higher cross-system efficiency, so evaluation must measure cycle-time and error-rate reduction, not just language scores.",[94,97,100],{"question":95,"answer":96},"What exactly is the difference between evaluating an agent’s final response and its trajectory?","Final-response evaluation judges whether the user’s goal was achieved correctly and helpfully — for example, whether a booking completed with zero errors or a balance update was accurate. Trajectory evaluation inspects the sequence of reasoning and tool calls: which tool was chosen, the parameters used, intermediate outputs, policy checks, and whether steps were efficient and safe. Together they reveal silent failures (a correct-looking answer produced via unsafe or incorrect steps) and support audits, compliance checks, and targeted remediation across multi-step workflows.",{"question":98,"answer":99},"How should teams instrument agents to support trajectory evaluation?","Teams must log every interaction-level artifact: prompts, model thoughts or step markers, tool invocation names and parameters, tool outputs, and full conversation histories. These logs should be structured, timestamped, and integrated into observability and incident pipelines so trajectory anomalies and safety violations generate alerts alongside latency and error metrics. Use automated eval tooling to score steps for validity, efficiency, and safety on each run.",{"question":101,"answer":102},"How do I align evaluation metrics with business value?","Translate language-oriented metrics into operational KPIs: measure cycle-time reduction, error-rate reduction versus legacy processes, workflow completion rates, and handoff success. Ensure dashboards report these business metrics alongside final-response and trajectory scores so executives see ROI (for example, “average days-to-close reduced” or “order error rate down X%”) rather than only language-quality numbers.",[104,112,119,126,131,135,140,145,151,158,164,170,176,184,190],{"id":105,"name":106,"type":107,"confidence":108,"wikipediaUrl":109,"slug":110,"mentionCount":111},"695e94e819d266277e14e030","AI agents","concept",0.99,"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAI_agent","695e94e819d266277e14e030-ai-agents",389,{"id":113,"name":114,"type":107,"confidence":115,"wikipediaUrl":116,"slug":117,"mentionCount":118},"69ebce4ae1ca17caac376494","Tier 2",0.92,"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FTier_2","69ebce4ae1ca17caac376494-tier-2",11,{"id":120,"name":121,"type":107,"confidence":122,"wikipediaUrl":123,"slug":124,"mentionCount":125},"699744639aa9beba177c5f9f","Silent Failure",0.9,null,"699744639aa9beba177c5f9f-silent-failure",2,{"id":127,"name":128,"type":107,"confidence":122,"wikipediaUrl":123,"slug":129,"mentionCount":130},"6a644855457d25046959d6da","Tier 0","6a644855457d25046959d6da-tier-0",1,{"id":132,"name":133,"type":107,"confidence":122,"wikipediaUrl":123,"slug":134,"mentionCount":130},"6a644855457d25046959d6d9","quality gate","6a644855457d25046959d6d9-quality-gate",{"id":136,"name":137,"type":107,"confidence":138,"wikipediaUrl":123,"slug":139,"mentionCount":130},"6a644855457d25046959d6d8","vibe-based testing",0.88,"6a644855457d25046959d6d8-vibe-based-testing",{"id":141,"name":142,"type":107,"confidence":122,"wikipediaUrl":143,"slug":144,"mentionCount":130},"6a644853457d25046959d6d3","final response","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FResponse_Boat_%E2%80%93_Medium","6a644853457d25046959d6d3-final-response",{"id":146,"name":147,"type":107,"confidence":148,"wikipediaUrl":149,"slug":150,"mentionCount":130},"6a644853457d25046959d6d4","trajectory",0.95,"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FTrajectory","6a644853457d25046959d6d4-trajectory",{"id":152,"name":153,"type":154,"confidence":108,"wikipediaUrl":155,"slug":156,"mentionCount":157},"695e3c6f19d266277e14dd48","OpenAI","organization","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenAI","695e3c6f19d266277e14dd48-openai",786,{"id":159,"name":160,"type":154,"confidence":130,"wikipediaUrl":161,"slug":162,"mentionCount":163},"6960e32919d266277e1504f6","Google","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FGoogle","6960e32919d266277e1504f6-google",248,{"id":165,"name":166,"type":154,"confidence":108,"wikipediaUrl":167,"slug":168,"mentionCount":169},"69600eff19d266277e14fa15","SAP","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSAP","69600eff19d266277e14fa15-sap",14,{"id":171,"name":172,"type":173,"confidence":115,"wikipediaUrl":123,"slug":174,"mentionCount":175},"6998c0419aa9beba177c778f","Tier 1","other","6998c0419aa9beba177c778f-tier-1",13,{"id":177,"name":178,"type":179,"confidence":180,"wikipediaUrl":181,"slug":182,"mentionCount":183},"6960e29a19d266277e150441","LangChain","product",0.98,"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLangChain","6960e29a19d266277e150441-langchain",196,{"id":185,"name":186,"type":179,"confidence":108,"wikipediaUrl":187,"slug":188,"mentionCount":189},"6a3dc9fac460e8b42cddc21c","Jalapeño","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FJalape%C3%B1o","6a3dc9fac460e8b42cddc21c-jalapeno",155,{"id":191,"name":192,"type":179,"confidence":148,"wikipediaUrl":193,"slug":194,"mentionCount":14},"6a5fb57e457d25046952b83e","Joule","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FJoule","6a5fb57e457d25046952b83e-joule",[196,203,210,218],{"id":197,"title":198,"slug":199,"excerpt":200,"category":11,"featuredImage":201,"publishedAt":202},"6a6457e207f0903672a88011","How NVIDIA’s Agentic and Physical AI Are Redefining Graphics and Simulation","how-nvidia-s-agentic-and-physical-ai-are-redefining-graphics-and-simulation","NVIDIA’s Vision: Agentic AI Meets Physical AI\n\n- Agentic AI:\n  - Systems that ingest diverse data, reason, plan multi‑step actions, and execute across tools\u002FAPIs, not just chat.[4]\n  - Deployed in log...","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1716967318503-05b7064afa41?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxudmlkaWElMjBhZ2VudGljJTIwcGh5c2ljYWwlMjBncmFwaGljc3xlbnwxfDB8fHwxNzg0OTYwOTkzfDA&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60","2026-07-25T06:38:03.167Z",{"id":204,"title":205,"slug":206,"excerpt":207,"category":11,"featuredImage":208,"publishedAt":209},"6a5fb2f9366a05b9f721d9f7","SAP Business AI Updates: How Joule Work and Enterprise AI Agents Redefine Digital Operations","sap-business-ai-updates-how-joule-work-and-enterprise-ai-agents-redefine-digital-operations","SAP’s latest Business AI updates move from “chat in a sidebar” to an AI execution layer that can run work across finance, HR, supply chain, and more. Joule Work, Joule Assistants, and Joule Agents sit...","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1759752394755-1241472b589d?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxzYXAlMjBidXNpbmVzc3xlbnwxfDB8fHwxNzg0NjU2NjMzfDA&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60","2026-07-21T18:10:05.488Z",{"id":211,"title":212,"slug":213,"excerpt":214,"category":215,"featuredImage":216,"publishedAt":217},"6a59ba596d00a851d4e57463","From Booth to Boardroom: How WAIC 2026 Exhibitors Can Showcase Production-Ready AI Systems","from-booth-to-boardroom-how-waic-2026-exhibitors-can-showcase-production-ready-ai-systems","WAIC 2026 lands squarely in what Stanford HAI calls the “evaluation era,” where the questions are “how well, at what cost, and for whom?” not “can AI do this?”[9]  \n\nBuyers and regulators will arrive...","safety","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1462826303086-329426d1aef5?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxib290aCUyMGJvYXJkcm9vbSUyMHdhaWMlMjAyMDI2fGVufDF8MHx8fDE3ODQyNjU2OTR8MA&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60","2026-07-17T05:21:33.547Z",{"id":219,"title":220,"slug":221,"excerpt":222,"category":11,"featuredImage":223,"publishedAt":224},"6a59a9c16d00a851d4e57290","Infrastructure and Supply-Chain Strain from Large Language Models","infrastructure-and-supply-chain-strain-from-large-language-models","The latest LLMs are no longer “just another cloud workload.”  \nEach new model family ramps compute, memory, and bandwidth needs, breaking old assumptions of near‑infinite elasticity.[2]\n\nGPT‑5.6 Sol m...","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1488272690691-2636704d6000?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxpbmZyYXN0cnVjdHVyZSUyMHN1cHBseSUyMGNoYWluJTIwc3RyYWlufGVufDF8MHx8fDE3ODQyNjEwNTd8MA&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60","2026-07-17T04:10:06.930Z",["Island",226],{"key":227,"params":228,"result":230},"ArticleBody_9GuCgpIm3yscbNavCVjN2ODs8bz5hCfFlsPtXBI94",{"props":229},"{\"articleId\":\"6a64468f07f0903672a87e6f\",\"linkColor\":\"red\"}",{"head":231},{}]