Agentic Voice-to-Form Automation

Filling out a form by hand is slow, repetitive, and prone to mistakes almost by design: the person already holds the information they need, but getting it into the form means locating each field, typing the right value into it, and repeating that once per field, every time. That friction is especially wasteful in cases where the information is something the user could just as easily say out loud rather than type: the bottleneck isn't a lack of information, it's the manual translation of what someone knows into a sequence of clicks and keystrokes.
The client needed a product built around that observation: instead of asking users to hunt through a form and type in each value, it would let them talk through their input naturally and have the system take care of getting each piece of information into the field where it belonged. That is a different problem than simple dictation, which just drops transcribed text wherever the cursor happens to sit. It requires understanding both what the user said and how the form is structured, and connecting the two automatically.
The genuinely difficult part of building that wasn't transcription itself, which is a comparatively well-understood problem. It was everything downstream of it. Speech is free-form and messy: people phrase the same fact in different ways, go out of order, add asides or corrections. Reliably turning that kind of unstructured speech into clean, structured data that maps onto a defined form schema, without asking the user to speak in some constrained, form-shaped way, was the real engineering problem.
Speech-to-Text Capture
A voice capture and speech-to-text flow that turns spoken input into accurate transcripts.
Agentic Field Extraction
An agentic AI pipeline extracts structured fields (not just keywords) aligned to the target form.
Automatic Form-Filling
Extracted values map to the correct fields automatically, ready to review and confirm.
Natural Speech Handling
Robust handling for corrections, asides, and out-of-order information in natural speech.
Review & Confirm UI
The user stays in control with a review step: AI output is a strong draft, never a silent commit.
The architecture was built around a handful of deliberate decisions, each targeting a specific failure mode in turning spoken language into structured form data.
- The pipeline was split into distinct stages (transcription, extraction, and field mapping) rather than one process that went straight from audio to filled-in fields. Keeping the stages separate meant each could be tuned and evaluated on its own terms, and any one of them, the transcription engine for example, could later be swapped out without touching the extraction or mapping logic built on top of it.
- Extraction relied on agentic LLM reasoning to pull structured data out of unstructured speech, rather than brittle regex or keyword matching. Pattern-matching only works within a narrow set of expected phrasings, and free-form speech doesn't stay inside one: people express the same fact many different ways and often out of order. Agentic reasoning let the system interpret what was actually said rather than search for fixed patterns, which was necessary for it to hold up against real speech rather than only the tidy cases it was built against.
- The mapping layer was built to be schema-driven, so the extraction layer targets a form definition instead of a hardcoded set of fields for one specific form. That decouples the reasoning logic from any single form: the same extraction approach can be pointed at a new schema and adapt to its fields, rather than requiring bespoke extraction code to be written for every new form the product needed to support.
- A review-and-confirm step kept the user in control instead of letting extracted values commit straight to the form. AI output is treated as a strong draft, not a final answer: the user can see where each value landed and correct it before anything is saved, so an occasional misread doesn't silently turn into bad data sitting in the record.
- The system eliminates manual data entry for the forms it covers: users speak, and the right values land in the right fields on their own, without anyone needing to hunt through the form and type each one in by hand.
- By treating capture and entry as the same spoken step rather than two separate ones, the system speeds up how quickly a form gets completed: the user talks once, and both providing the information and getting it into the form happen together instead of one after the other.
- Because the pipeline is staged and the mapping layer is schema-driven, the result is a reusable agentic foundation rather than a solution built for a single form. It can be extended to new forms by pointing it at a new schema, and to other data-extraction use cases that share the same underlying problem of turning unstructured input into structured output.



Discovery
Understand the problem, the users, and the real constraints before any design work starts.
Design
Architecture and UX decisions made deliberately, before a line of production code is written.
Build
Iterative development with regular check-ins, so direction can be corrected early and cheaply.
Test
QA and hardening against real-world edge cases before anything reaches production.
Ship & Support
Launch, then stay involved: monitoring, fixes, and iteration as the product keeps growing.

