Contact me
arrow_back

Data Studio · Innovaccer · 2021 – 2025

Only our experts could run it. We turned it into something customers run themselves.

Healthcare data used to take six months to go live. We brought it down to a day for some sources, two weeks typical, and cut the accuracy gap by 60%. I was a designer on this from 2021 to 2025.

The problem

It worked because our experts ran it by hand

When I joined in 2021, Innovaccer’s data capability was already the backbone of a three billion dollar company. Every insight the platform sold ran on data our teams had extracted, mapped, and modelled. There was just one catch: only our own experts could drive it. The company was going from 1 to 10, and the question handed to us was never “fix this.” It was “how does this scale without an expert in the room?”

Everything worked because our people were exceptional at operating it. Extraction ran over VPN and shell, driven by engineers who knew every customer’s setup by heart. The ETL orchestrator did its job as a scheduler while the real thinking happened in the tools our experts preferred. The data model lived in the heads of the two modellers who built it. That is what winning looks like before you turn it into a product. It also took six months per customer, and six months does not multiply.

Customer onboarding, touchpoint by touchpoint2021 baseline · 50-80 sourcesRebuilt from memory
01
Kickoff + project scoping
attacked by standardization · ch1
8-10 days
02
Environment setup
attacked by warehouse · ch1
~2 weeks
03
Access, walkthrough, extraction install
replaced VPN + shell · attacked by RDE · ch1
~1 week
04
Data model mapping
the two-modellers bottleneck · attacked by modelling · coda
from memory: unclear
05
ETL / pipelines
attacked by copilot · ch3
5-6 weeks
06
Ontology mapping
~40% of codes bounced to customers · attacked by semantics · ch2
3-4 weeks
07
Analytics activation
attacked by quality rules · ch2/3
3-4 weeks
08
Analytics review
from memory: unclear
09
Visualization layer / dashboards
including custom builds
5-6 weeks
10
App-sync layer
from memory: unclear
Total ≈ 5-6 monthsJan 2025: ~2 weeks typical · under 24 hours for some sources

The journey map we kept coming back to. It measured the first chapter, picked the second, and pointed to the third. The original was lost; this is rebuilt from memory.

The users

The expert in the room, and the team without one

For years the expert was always in the room. Our own engineers knew each customer’s setup by heart, two modellers carried the data model in their heads, and the activation team drove every onboarding by hand. That is who the product really was, and it worked.

The bet was to move all of it to the other side of the table: the customer’s own data team, running the same platform with nobody from Innovaccer in the room. Every decision that follows is really one decision, asked over and over. What has to disappear for that to be safe?

The bet

Make it feel like a consumer app

The product was as data-dense as they come, and the bet sounded backwards: remove so much complexity that using it would feel like a consumer app. That is a bet about how the product behaves, not about what the screens show. Take the hard decision away from the user entirely, and the interface stops being the hard part. Every chapter that follows is a different version of the same move.

ETLCustomer-run
Patient · source mappingscoverage 96%
adt_hl7PID-5 · namepatients.namemapped
adt_hl7PID-7 · dobpatients.birth_datemapped
eligibility_834member_idpatients.member_idmapped
crm_export.csvpreferred_langpatients.languageproposed
AI proposes, the quality rules judge, failures go to a human.

The bet, working: source feeds map into the model, proposals clear the backlog, and the hard decision never reaches the user. Representative data.

Decision 1 · Warehouse · 2021–2022

We went all in on the warehouse

My first brief: make it so we never have to shell into a customer’s machine again. I thought we could email an installer with instructions and be done. Then I learned how US care networks and their facilities actually work. They wanted us to walk them through the product before anything touched their machines. That was my first real lesson in how seriously this industry takes HIPAA.

Then came the fork: more connector types, or the warehouse. The answer lived in a Google Sheet the activation team kept, one row per customer. I collated it and cleaned it myself: 90% of our use-cases ran on SFTP and on-prem extraction. That settled the fork. What it did not settle was how to organize the warehouse, and that became the longest argument of my time at Innovaccer.

Read the scene: the design sprint

Reconstructed from memory. October 2022. Five ways to organize a warehouse, and every one of them was partially right.

Design wanted the simplest structure. Engineering wanted the fastest queries. Product wanted the lowest cloud cost. Three fair positions, three different answers, and months of data gathering that made every case stronger without settling anything. My design manager finally put senior directors and the working team in one room for a design sprint.

My PM and I presented the interview data and one reframe: take warehouse access and navigation away from users entirely. Show everything through the platform. Reduce the complexity until nobody cares where the data lives.

Once the structure stopped mattering to users, it stopped being a design argument at all. It became a pure engineering efficiency problem, and my EM proposed the final structure from that angle the same day.

Nobody won that argument. We made it unnecessary.

Decision 2 · Quality · 2023

Cleaning up the codes

With data flowing, the journey map pointed at the slowest remaining step: roughly 40% of healthcare codes arrived non-standard and bounced back to customers to resolve by hand. We bought a mapping vendor. There was no argument to lose this time. Mapping medical code systems was not the business we were in, and we needed it to go well. Two years later I sat on the other side of the same principle, arguing to build: that argument opens the AI Studio story.

The result: the gap between our calculated outcomes and what payers reported dropped by 60%. Patient identity ran as this chapter’s twin: same shape, same boundary. I designed the reporting pipelines and interactions; other designers owned the final customer-facing screens. My wrong version taught me the most: a big end to end design I drew before the MVP had met a single user. I have not designed the cathedral before the chapel since.

Semantic qualityReportingRebuilt
Codes arriving non-standard, by category
Labs · LOINC58%
Procedures · CPT/HCPCS46%
Medications · NDC/RxNorm39%
Diagnoses · ICD-1031%
Encounters · ADT14%
≈40% overall · every non-standard code routes back to the source that sent it
-60%
accuracy gap vs payer-reported outcomes, after the mapping service
Standard mapping
your team researches raw codes and shares mappings back · 3-4 weeks
Mapping service · opt-in
vendor-powered, cost passed through · proposed mappings arrive for your review in days

My slice, exactly: the quality reporting and the routing of non-standard codes back to their sources. Another designer owned the review screens. Representative data, iterated past what shipped. Opt in to watch the service land.

Decision 3 · Co-pilot · 2023–2025

We gave AI the scripts and kept the decisions

I inherited the Data Co-pilot as a form-based wizard for one data type when its designer left. GPT was just catching wind and we were deliberately conservative: we assigned AI the scripting tasks and kept the decisions. The evals had a seed here too. My PM had introduced data quality rules that were, in hindsight, evals for data. AI proposed mappings, the rules judged them, failures went to a human. We called it human in the loop back then. In 2025 we would call the same instinct something else.

Data Co-pilot · SQL

Synthetic sandbox: patients, encounters, claims, providers, coverage, labs. The AI writes scripts; the decisions stay yours.

The instinct, working: a real SQL editor over sandbox data, running in your browser. Ask in plain English and a live model writes the query; a gate of fixed rules keeps it read-only and inside the schema. Scripts from the AI, decisions with you.

The receipt

We hit the target set in 2021

By January 2025 the co-pilot had come together as one product, and the numbers were in. Data went live in under 24 hours for some sources, two weeks typical, which we reported conservatively as an 85% cut from six months. The accuracy gap against payer-reported numbers had dropped 60%. And the backbone only our experts could run had become a family of products, modelling and cataloguing and governance, that customers operate themselves.

The win bought leadership confidence, and confidence became scale. I ran the new data products from March 2025 until the offload that August, when AI Studio took my full bandwidth. The backbone I was leaving would feed the company’s next bet.

Data model
patients
idint
nametext
birth_datedate
languagetext
member_idtext
encounters
idint
patient_idpatients.idint
classtext
departmenttext
start_tsdate
claims
idint
patient_idpatients.idint
provider_npiproviders.npitext
billed_amountreal
paid_amountreal
statustext
providers
npitext
nametext
specialtytext
coverage
idint
patient_idpatients.idint
plantext
end_datedate
labs
idint
patient_idpatients.idint
loinctext
valuereal
observed_atdate

Drag the tables, pinch or ctrl-scroll to zoom, scroll to pan. Primary keys in accent; every arrow is a foreign key. The seeded six are what the SQL studio queries; your extensions live on this canvas.

Where it all ended up: the model two modellers carried in their heads, now a canvas you can extend by hand or by asking. A demo, not a screenshot.

Closing

The expertise moved into the product

My design decisions were about how the product behaved, more than what the screens showed.

A customer runs the platform today without ever meeting the expert who used to be in the room; that is where the expertise went. None of these screens are product screenshots. Everything here is rebuilt in my own design library, iterated past what shipped.

Sequel: AI Studio, the bet this platform fed

Researched, ideated, designed, and shipped with love and care. View how · Resume