Skip to content
ai-agent.review — AI Agent Reviews, home
← All news

Research

G2’s CX agent evaluation beta separates task results from buyer reviews

Vendor-reportedChecked: 2026-10-06[g2-launch][g2-method]

Editorial attribution: Gosia Madej

Prepared with AI assistance. Editorial attribution does not imply personal product testing or manual review of every story.

Published: · Event date:

The September 14 announcement introduced a public beta. Checked October 7, its general CX leaderboard shows Tidio and Retell tied at 200 points.

Editorial assessmentChecked: [g2-launch][g2-board]

Two equal stacks of completed task tiles stand beside a magnifying glass and a separate customer feedback bubble.
Original editorial illustration of task evidence and separate buyer feedback; not a G2 interface or a chart of product scores. · AI Agent ReviewsImage use and permissions

G2 described its first customer-experience agent leaderboard as live in public beta on September 14. Documentation dated September 8 already described the program, so September 14 is the dated announcement used here, not a verified first-access date. This is a review of the September release, not a launch today.

Vendor-reportedChecked: [g2-launch][g2-docs]

The general leaderboard currently lists nine evaluated agents. Tidio and Retell each have 200 Overall points, followed by LiveAgent at 195. Although the display places Tidio first and Retell second, their published totals are equal. G2 lists accuracy/compliance at 87%/92% for Tidio and 86%/90% for Retell; these are separate dimensions, not evidence of universal superiority.

Independent reportChecked: [g2-board]

G2’s September 22 scoring note describes 46 support tasks in a simulated business with 38 tools. Each passed standard task earns five points, making 230 the current maximum; 200 therefore represents 40 passed tasks, not 200%. Failed or missing tasks earn zero. G2 says each task currently runs once, with repeated trials planned.

Vendor-reportedChecked: [g2-scoring][g2-method]

Accuracy measures evidential support for claims, compliance checks the supplied policy, and relevance measures focus on the customer’s need; G2 uses LLM judges for these. Completeness checks expected system outcomes deterministically. Supporting dimensions are normalized percentages, separate from the Overall points. Buyer ratings are a third signal and do not change the evaluation score.

Vendor-reportedChecked: [g2-method]

The release has material comparison limits. G2’s announcement says its beta setup favors headless agents and does not yet fully represent established, configured workflows. The scoring note describes a separate retail task set for ecommerce agents. G2 documentation says vendor evaluations are free during beta, with paid evaluations planned.

Vendor-reportedChecked: [g2-launch][g2-scoring][g2-docs]

Checked against sources

AI-assisted source check

AI-assisted checking read G2’s announcement, documentation, leaderboard, scoring explanation, technical draft and Tidio result page. G2 is the program operator and a third-party evaluator of the listed products. These pages are one reporting origin; we did not rerun the benchmark or independently audit its grading.

Editorial assessmentChecked: [g2-launch][g2-board][g2-method][g2-scoring][g2-docs][g2-tidio][g2-draft]

Dates describe different things: Tidio’s result page dates its run to August 6, while the announcement and scoring explanation appeared in September. The linked August 20 technical PDF is marked draft and describes score-combination modes that differ from the newer pass/fail explanation. We use the newer published explanation for current points and flag the documentation mismatch rather than silently merging formulas.

Editorial assessmentChecked: [g2-tidio][g2-scoring][g2-draft]

Source evidence

Original source excerpt. This shows what the source says; it does not establish product performance.

Our first leaderboard for Customer Experience AI Agents is live in beta.

Read original source [g2-launch]

What changes for users

Use this as a shortlist input, then inspect task traces and confirm the evaluated configuration, date and use case match your planned deployment. A passed task can still contain a policy mistake: the overall total and supporting dimensions answer different questions. Require repeated tests on your own policies before inferring production reliability, cost savings or ROI; this snapshot establishes none of those outcomes.

Editorial assessmentChecked: [g2-method][g2-scoring]

Sources