Harshith
Beecha.
← Back to the story
Chapter 4 of 43 min skim · expand any section for the full story
Built on the conversation object from Chapter 3.

Cutting enterprise agent load by 40% with AI voice

Turns out, AI can handle that call.

AI / VoiceEnterpriseIVRHealthcare
40%
Drop in agent load
10+
Enterprise accounts
2
Verticals: healthcare & recruiting
The Conversive agent inbox with a completed voice call: the call recording in the conversation thread, an auto-generated call summary with action items, and a timestamped live transcript panel alongside it

A completed voice call in the agent inbox: the recording, an auto-generated summary with action items, and the timestamped transcript.

By 2024, voice AI crossed from experiment to enterprise-ready. Conversive acquired Voxgenie and greenlit a voice product for IVR modernisation and inbound/outbound call handling in healthcare and recruitment. I owned the research, use-case strategy, and the evaluation framework that made it enterprise-grade, landing a directional 40% reduction in agent handling time across the first 10+ rollout accounts.

Context

CompanyConversive by SMS-Magic
TimelineJuly 2025 (research) – present (rollout ongoing)
RoleAssociate PM (owned the voice AI product workstream)
VerticalsHealthcare, Recruitment
Team
EngineeringSalesCS

For most of its history Conversive was text-first: SMS, WhatsApp, Email. Voice was the gap, and by 2024 enterprise buyers were asking questions we couldn’t answer.

I led market and competitive research across the voice AI landscape, defined the use-case strategy, wrote the quality-requirements PRD, and built the AI agent evaluation framework covering scenarios, testcases, and evaluators for systematic quality testing.

Engineering owned the Voxgenie integration architecture and model selection; sales and CS owned the account rollout relationships. I owned use-case strategy, the quality bar, and the evaluation framework.

The Problem

Healthcare and recruitment both run high-volume, high-stakes phone calls that are expensive to staff and hard to scale, from candidate screening to appointment reminders to patient follow-ups, while legacy IVR systems frustrated callers with rigid touch-tone menus.

After-hours calls were going to voicemail. Missed appointments, lost revenue.
Legacy IVR

Rigid touch-tone menus that frustrate callers and can’t handle natural language.

Voice AI

Understands natural speech, handles structured calls end-to-end, and hands off cleanly when it can’t.

Discovery

The market had three layers, none complete

Model companies, horizontal platforms, and vertical specialists, but no CRM had a clean, native voice AI story.

Build on models, not platforms

Existing platforms meant faster launch but commoditised pricing; model platforms meant more work but a real product layer we could own.

Enterprises wanted a wedge, not a leap

After-hours overflow, structured outbound, and IVR modernisation were the low-risk entry points, not full call automation.

Source: landscape mapping across model companies, horizontal platforms, and vertical specialists · direct customer conversations

Salesforce had Agentforce but required Amazon Connect or ISV partners for voice, plus significant customisation. HubSpot supported voice via Breeze AI but had no automation, so customers needed a second RetellAI subscription. Zoho Voice had simple workflows but no AI capabilities. None of them had a clean, native voice AI story, which was a clear opening for Conversive to be the voice layer CRM-native customers didn’t have to stitch together themselves.

Key Decisions

01

Target constrained, high-volume calls first

Structured, measurable, high-volume. Recruitment screening and healthcare appointments fit; escalations and negotiations don’t, yet.

Voice AI works best when calls have a clear structure and a measurable outcome. Recruitment screening and healthcare appointment flows are exactly that: a defined set of questions, a binary outcome, and enough volume that the economics work easily. We explicitly deprioritised open-ended, high-complexity calls (escalations, complaints, negotiations) for Phase 1, because the risk/trust bar was too high and the success criteria too fuzzy.
02

Quality as a prerequisite, not a feature

Seven foundational requirements had to be met before any enterprise account, none of them optional.

Shipping a voice product that felt robotic or unreliable would be worse than not shipping, especially in healthcare and recruitment, where a bad call directly damages trust. I defined seven foundational requirements that had to be met before any enterprise account: immediate agent response (0.5–2s), accurate live transcription during interruptions, clean end-of-turn detection, natural ambient sound, backchanneling cues, p95 latency under 2 seconds, and stable background-noise handling. None of these are features. They’re table stakes.
03

Build a systematic evaluation framework

Manual testing doesn’t scale. Four independent layers catch failures a single pass/fail check would miss.

I built a structured evaluation pipeline with four layers: capability extraction (testable rules from the agent’s system prompt), scenario generation (happy path, hesitant users, edge cases, knowledge boundary tests), testcase generation (3–5 realistic multi-turn conversations per scenario), and three independent evaluators for scenario outcome, capability assertions, and KB grounding. An agent can pass its scenario outcome but still violate a capability rule; evaluating them independently surfaces different classes of failure.

How It Works

A structured pipeline tests agent quality systematically instead of ad hoc, four layers deep.

01

Capability extraction

Testable rules pulled straight from the agent’s system prompt.

02

Scenario generation

Happy path, hesitant users, edge cases, knowledge boundary tests.

03

Testcase generation

3–5 realistic multi-turn conversations per scenario.

04

Three independent evaluators

Scenario outcome, capability assertions, and KB grounding, scored separately.

What We Built

IVR modernisation. Natural language replaces touch-tone menus; the agent understands intent and routes accordingly.

Automated inbound handling. AI collects info, routes to humans when needed, or resolves independently. No more voicemail after hours.

Outbound call automation. Structured flows for candidate screening and appointment reminders, logged straight to the CRM.

Recruitment & healthcare use cases. Candidate screening and interview scheduling; appointment reminders and front-desk overflow.

AI agent quality framework. Systematic evaluation pipeline for repeatable quality measurement across agent versions.

The voice agent analytics screen: calls received, resolved by agent, transferred to human, and abandoned in queue, with a call progress funnel, an outcome breakdown, and a per-intent containment table

The analytics surface teams read after rollout: containment, transfer rate, and where handle time goes, broken down by intent. Sample data.

Outcome

A directional 40% reduction in agent handling time across 10+ enterprise accounts in healthcare and recruitment. That is a blended figure, not a controlled A/B test. The structured, repeatable calls we targeted first were exactly where AI handled the most volume with the least human escalation.

Beyond the metric
Recruiters and healthcare staff shifted to higher-value work
A repeatable eval pipeline now catches regressions before release

What I'd Do Differently

Should’ve defined success per use case, not blended

A portfolio metric hides which use cases are actually working.

"40% reduction in agent handling time" is a portfolio metric. It doesn’t tell you which use cases are working and which aren’t. Recruitment screening and healthcare appointment reminders have different success shapes: one is about qualification accuracy, the other about show rates. I’d instrument per-use-case from the start so we could double down on what’s working and fix what isn’t, rather than optimising for the blended number.

Underinvested in the handoff experience

When a call escalates to a human, losing context breaks trust fast.

When a voice AI call escalates to a human agent, the transition is a critical moment. If the agent doesn’t pass context cleanly, meaning what was asked, what was said, and where in the flow the caller is, the customer has to repeat themselves and trust collapses. We solved for it, but I underestimated how much friction that handoff could still create in real deployments. I’d treat escalation quality as a first-class product requirement from day one, not a polish task.

That's the whole story. Back to the top ↑