How AI-Powered OHIP Billing Software Affects Time, Revenue, and Workflow for Ontario Family Physicians
OntarioMD was independently contracted to evaluate Quip's AI billing solution as part of a formal pilot project, led by Dr. Yoobin Lee, Dr. Arun Radhakrishnan, and Dr. Zack van Allen. Their analysis compared physician-reported outcomes and actual Remittance Advice (RA) data from before and after Quip adoption. Below is a summary of what OntarioMD found, including where the data was strong, where it was inconclusive, and where it simply wasn't enough to draw a conclusion.
Methodology
OntarioMD's evaluation drew on two data sources:
- A pre/post physician survey, with 18 respondents providing usable data.
- An analysis of Remittance Advices (RAs) from 12 respondents.
Data was cleaned and modelled in Excel Power Query and Power BI, with descriptive statistics (median, mean, frequencies) alongside nonparametric paired-sample tests (Related Samples Marginal Homogeneity) and paired-samples t-tests at an alpha of .05.
Physicians were compared to themselves, before and after using Quip, rather than to a separate control group - a reasonable approach for a pilot of this scale, though the sample size (12–18 respondents) means some results should be read as directional rather than definitive.
The two findings with statistical significance
Two outcomes cleared OntarioMD's significance threshold (p<.05) with minimal ambiguity, meaning the pattern was consistent enough across respondents that chance is an unlikely explanation.
Billing time per patient encounter decreased
Physicians were asked how much time they spent on billing-related tasks per encounter, including code lookup, documentation review, claim submission, and after-hours review. The shift was substantial: the proportion of physicians spending less than 5 seconds per encounter on billing nearly doubled, from 16.7% pre-Quip to 33.3% post-Quip, while the "30 seconds to 1 minute" category dropped from 33.3% to 11.1%. OntarioMD noted this as one of the clearest, least ambiguous findings in the entire study, likely because it maps most directly onto what an automated billing tool is actually designed to change.

Average daily revenue increased
Pulled directly from RA data rather than self-report, average revenue per day rose from a pre-Quip median in the $350–450 range to a post-Quip median noticeably higher: a median increase of $89 and a mean increase of $196 per day, with paired-samples t-tests confirming statistical significance and a moderate effect size. Monthly revenue from clinic work showed a similar upward trend (median increase of $1,116, mean increase of $2,000), though this result landed just short of statistical significance (a "near significant" finding worth watching as the sample grows, not yet one to treat as conclusive).

Revenue gains varied by gender
A breakdown of the same revenue difference data by physician gender points to a larger post-Quip gain among female participants than male participants:
| Gender | N | Mean Revenue Diff | Std. Deviation | N (Avg. Revenue Diff) | Mean Avg. Revenue Diff | Std. Deviation |
|---|---|---|---|---|---|---|
| Female | 7 | $2,828.57 | $5,101.62 | 7 | $217.21 | $330.20 |
| Male | 5 | $839.80 | $438.65 | 5 | $166.59 | $289.96 |
| Total | 12 | $1,999.92 | $3,913.43 | 12 | $196.12 | $301.20 |
Both the overall revenue difference and the average daily revenue difference were higher among female physicians in this sample. It's worth reading this as a descriptive pattern rather than a confirmed effect: the table reports means and standard deviations by group, not a formal statistical test between the two groups, and splitting an already-small sample (N=12) into subgroups of 7 and 5 leaves little room to distinguish a real effect from noise. The female group's standard deviation ($5,101.62) is also large relative to its mean, suggesting the average is being pulled up by one or two higher-performing outliers, consistent with the outliers OntarioMD flagged in its original scatterplot analysis.
Whether this reflects differences in prior billing practices, panel composition, or something else specific to this small sample isn't something this pilot can answer, but it's a pattern worth tracking in the future.
Directional findings: promising, but not yet statistically confirmed
Several other measures moved in a positive direction but either lacked the statistical power to confirm significance or showed enough variation across respondents that OntarioMD flagged the results as ambiguous rather than clear-cut:
- Discovery of new billing codes and premium opportunities trended toward more frequent discoveries post-Quip, though zero cell counts in parts of the data prevented formal significance testing.
- Perceived accuracy of billing submissions shifted meaningfully, with the most common response moving from the 50–70% range pre-Quip to the >90% range post-Quip. Again, significance testing wasn't possible here due to data structure, so this should be read as a trend rather than a proven effect.

- EMR-related frustration showed a near-significant decrease (p=.074), with the "Agree that my EMR adds to my frustration" response dropping from 50% to 11.1% pre- to post-Quip. Just outside the conventional significance threshold, but a meaningful shift worth monitoring in a larger follow-up.
- Feeling undistracted by billing concerns during patient appointments showed a statistically significant shift in distribution (p<.05), but the pattern itself was mixed: "strongly agree" responses increased sixfold (5.6% to 33.3%) while plain "agree" responses fell (77.8% to 38.9%) and "disagree" responses also rose slightly (5.6% to 16.7%). OntarioMD's read is that Quip affected different physicians differently here, likely reflecting clinic- or clinician-level factors the study wasn't designed to capture.
What showed no significant change
Several outcomes showed no statistically significant pre/post difference, and in some cases contradictory movement that OntarioMD attributes to factors outside the study's scope (individual clinic context, workload, staffing) rather than to Quip itself:
- Same-day billing completion rates
- Work-related stress and general job satisfaction
- Perceived administrative burden on wellbeing
- Sense of control over workload
- Care team efficiency and alignment with clinic leadership values
What this means if you're considering AI-powered OHIP billing automation
The strongest, most defensible takeaway from an independent pilot evaluation of this kind is narrow but real: physicians using Quip spent measurably less time on billing per patient encounter, and their actual OHIP remittance data showed a statistically significant increase in average daily revenue. Many other factors in the study, such as code discovery, perceived accuracy, distraction during appointments, EMR frustration, moved in a positive direction without yet reaching the bar of statistical certainty. OntarioMD's own postamble captures this balance well: clear evidence of time savings and revenue gains, alongside acknowledged limits in statistical power and some genuine ambiguity in the softer, more subjective measures.
This summary reflects an independent pilot evaluation conducted by OntarioMD, a wholly owned subsidiary of the Ontario Medical Association, under a Pilot Project Agreement with Quip Medical Inc. Quip Medical builds AI-powered OHIP billing optimization for Ontario physicians.