Evaluation worksheet
How to evaluate AI patient intake before a patient rollout
Evaluate the conversation and report using your own questionnaires before expanding to patients. Agree what a useful result looks like, then measure the full workflow, including review, corrections and support.
- See Emily liveReview the conversation and notes
- Try your customized agentUse your questionnaires in the portal
- Try the patient workflowAgree the workflow and measure results
What does a useful intake trial need to prove?
A trial should establish whether the conversation collects the required history, whether the notes faithfully represent it and whether the full workflow reduces work for your practice. A fluent conversation or a well-formatted report alone cannot answer all three questions.
AHRQ’s health IT evaluation resources encourage choosing measures that match the purpose of the implementation. [1] Write the primary question before the trial begins: for example, “Does this reduce total physician history and documentation time for our selected visit types while maintaining an acceptable clinical handoff?”
Agree the population, current process, reviewer, timing method, patient support and decision criteria. Keep new visits and follow-ups distinguishable. Record the questionnaire and note-template version so later changes do not silently alter what you are evaluating.
Stage 1: test the conversation and report together
After a live demo, evaluate an agent customized with your questionnaires and physician note preferences. Use sample inputs that reflect the variety of your routine work, including both straightforward and difficult-to-summarize accounts.
| Case to test | What the reviewer should look for |
|---|---|
| A clear new concern | Timeline, relevant context and patient priority retained |
| A follow-up with partial improvement | Both improvement and remaining difficulty captured |
| An uncertain date or medication detail | Uncertainty preserved, with no invented answer |
| A patient corrects an earlier statement | The correction reflected without silently keeping the wrong fact |
| A question is skipped or the session ends | Incomplete information marked clearly |
| A relevant movement is shown | Observation separated from what the patient reports and from clinical interpretation |
| A person needs device or language support | An appropriate assistance route and visible completion status |
How do you assess note quality beyond an accuracy percentage?
Define what a correct and complete history means for each case before reviewing the output. Use the patient’s actual account and the physician’s required elements as the reference. A single overall score can hide a clinically important omission or an unsupported statement.
| Dimension | Example of a problem to record |
|---|---|
| Factual fidelity | Wrong side, date, sequence or treatment response |
| Completeness | An important concern is missing |
| Unsupported content | The note adds a symptom denial or finding the patient did not provide |
| Uncertainty and source | Patient report is presented as a confirmed clinical finding |
| Clinical usability | Relevant information is buried in repetitive text |
| Format and handoff | The note does not fit the agreed physician or EHR format |
Classify issues by clinical importance and by the work required to correct them. Have the clinical reviewer resolve ambiguous cases. A useful evaluation records both errors and successful preservation of important details; it should not reward longer notes simply for containing more text.
Stage 2: agree the workflow before patient use
Identify who sends the invitation, who supports patients, who reviews incomplete or concerning responses, and how notes reach the clinical record. Review patient information, access arrangements and applicable agreements before live use. The practice’s clinical protocol should determine how urgent concerns are handled.
Confirm the exact EHR connection separately from the conversation test. FaceMed.ai supports integration with multiple mainstream EHR systems. Connection scope, note delivery and timing are agreed for your EHR and workflow. Without automatic integration, launch can take as little as a few hours once questionnaires and note preferences are ready. EHR integration timing depends on the system and agreed scope.
AHRQ’s workflow toolkit is useful here because it treats clinical and administrative work as connected parts of implementation. [2] Write down the exception route as well as the ideal route.
Stage 3: measure time with consistent definitions
Compare similar encounter types in the current and trial workflows. Count all physician time used for history and documentation, including report review, clarification, correction and final record work. Keep staff time and patient effort in their own columns.
- Direct observation: useful for timing tasks, but define start and stop points and avoid double-counting overlapping work.
- EHR activity logs: useful for recorded application activity, but may miss work in another tool or conversation time.
- Physician estimates: useful for perceived workload, but label them as estimates rather than timed observations.
Report the mean and the median when the sample supports both, along with the spread and number of encounters. Include incomplete intakes and visits where a report was available but not used. Log changes in staffing, case mix and integration during the trial.
Printable intake evaluation worksheet
Keep patient identifiers out of this worksheet. Specify the eligible population and count exclusions with their reasons. Blank cells mean not measured, rather than zero.
| Measure | Current workflow | Trial workflow |
|---|---|---|
| Visit type, date range and number of eligible encounters | ||
| Invited / started / completed / reviewed before visit | ||
| Not completed, not used or needing help, with reasons | ||
| Physician minutes: history and all documentation tasks | ||
| Review, correction and transfer minutes not already counted | ||
| Staff minutes: preparation, reminders and support | ||
| Patient time, accessibility and experience feedback | ||
| Important omissions, inaccuracies and unsupported statements | ||
| On-time starts and after-hours work | ||
| Recurring costs and one-time setup costs, separately |
Print this page to use the worksheet. Count every task once. For completion rate, divide completed intakes by invited patients and report the invited share of all eligible patients separately.
When should you expand, revise or stop?
Agree the thresholds with the clinical and operational owners before looking at the result. Expand when notes meet the agreed quality standard, the patient pathway works and the measured benefit justifies the ongoing cost. Revise when a fixable handoff or questionnaire issue consumes the benefit. Pause when an important quality or workflow issue is unresolved.
AHRQ’s Plan–Do–Study–Act guidance supports small tests followed by review and adjustment. [3] Do not expand solely because users liked the demo or because one physician had an unusually strong result.
- How many encounters are enough for a trial?
- There is no universal number. Include enough routine variation to assess the workflow and report the actual sample. A trial can establish local feasibility; a precise generalizable performance claim needs a study designed for that purpose.
- Can we evaluate Emily before automatic EHR integration?
- Yes. A customized portal evaluation can test the conversation, questionnaires and note format first. Keep integration work and record-transfer time visible when moving to the complete patient workflow.
Sources and further reading
Published by FaceMed.ai. For practice planning and evaluation; clinical decisions remain with your care team.