01
About
Industry
Pediatric Health
Product Type
Enterprise AI Transformation, Voice AI Agent
Services
02
Objectives
Automate initial call handling and multilingual patient data intake through a Wildix Voice Agent to reduce front-desk congestion and expedite lead collection.
03
Challenges
Designing a conversational voice agent for a pediatric therapy provider required an especially careful approach to communication. Because the agent would interact with parents and guardians in situations that could involve concern, urgency, and emotionally sensitive topics, it needed to respond with empathy while maintaining strict operational boundaries. A central challenge was ensuring the agent would never provide diagnostic guidance or clinical interpretations, even when directly prompted, and would instead guide callers toward scheduling an appointment with the company.
Another important challenge was preparing the knowledge base that would support accurate and safe responses. The information needed for the agent was distributed across the company’s existing ecosystem and was not immediately structured for conversational use. Extracting, cleaning, and organizing that content into a format suitable for a voice-based GPT agent required both precision and contextual understanding, particularly given the sensitivity of pediatric healthcare interactions.
The project also required a fully multilingual experience in English and Spanish. This introduced additional complexity beyond direct translation, as the agent had to preserve the same tone, safeguards, and workflow logic across both languages. Ensuring consistent performance in bilingual conversations was essential to delivering a reliable front-line experience for a diverse caller base.
In addition, the voice agent was expected to perform conversational intake when a caller expressed interest in becoming a patient. While this capability was central to reducing front-desk workload and accelerating lead capture, collecting contact information accurately over the phone presented a practical challenge. Names, phone numbers, email addresses, and other intake details are often difficult to capture correctly in a live voice interaction, especially when speech patterns, accents, background noise, or spelling variations are involved.
Finally, validating the performance of the agent at scale posed its own challenge. Manual testing of multi-turn voice conversations is inherently time-consuming and difficult to standardize, particularly when the system must be evaluated across multiple languages, edge cases, and adversarial scenarios. A more rigorous and repeatable testing approach was necessary to ensure the agent could perform reliably before deployment.
04
Solutions
To address the sensitivity of pediatric healthcare interactions, the agent was shaped through extensive prompt engineering, carefully defined conversational guardrails, and repeated validation cycles. These measures ensured that the voice agent could handle emotionally charged or medically oriented questions in a safe and professional manner, declining to provide diagnoses or clinical opinions while steering parents toward scheduling an appointment. The same framework was used to prevent the agent from answering billing and insurance questions, allowing it to remain within its intended scope and reinforcing trust in every interaction.
To support this behavior with accurate organizational knowledge, the required information was extracted and refined from the company’s broader ecosystem through a human-assisted scraping process supported by large language models. This approach made it possible to transform fragmented internal content into a usable knowledge layer for the agent while preserving quality and relevance.
For multilingual delivery, the team combined prompt engineering with the appropriate Wildix configuration to ensure the agent could operate effectively in both English and Spanish. Dedicated test cases were created for each language so that tone, routing logic, safety controls, and data collection flows could be validated independently. This made it possible to deliver a bilingual experience that remained consistent in functionality and user experience.
To improve conversational intake, the team iterated heavily on the way the agent asked for and confirmed caller information. Prompt design was optimized to make spoken data collection more natural, accurate, and resilient during live phone interactions. Through repeated testing and refinement, the agent became more effective at capturing lead details for follow-up while preserving a smooth caller experience.
To overcome the limitations of manual QA, testing was heavily automated in a sandbox environment using the DeepEval framework. Multi-turn conversations were simulated with an LLM User Simulator, while an LLM-as-a-Judge supported the review of conversation quality and policy adherence. The system was tested against synthetic scenarios derived from real inbound call patterns, along with adversarial cases designed to probe weaknesses and reinforce reliability. Although the process remained under human supervision, this structured and automated methodology made testing significantly more efficient, scalable, and rigorous.
We'd love to hear from you!
Drop us a message, and let’s start creating something amazing together.


