A Decade of Building with AI

The Healthcare Antidote in the Making

It’s late 2015 and after watching six figures of my own money vanish in a failed AI startup, I was on the verge of giving up on my dream of ever becoming an entrepreneur.

My nervous system went for a toss and I was riddled with physical pain, this was my first entry point into the healthcare system.

By now, AI was alive but this time around it wasn’t AI that sparked the journey, it was frustration. A deep, firsthand encounter with how broken the healthcare system was. How deeply flawed the process was from diagnosis to investigations to treatment. It felt like trial and error conducted by a human in a white coat.

There I was, impatient and irritated in a doctor’s office, picking apart the flaws of a failing system.What I didn’t realise then was that my irritation was quietly shaping the seed of my next venture.

Within months, I had co-founders, a committed team, and investors eager to reimagine healthcare at its roots. But as the work unfolded, one truth became undeniable: this system was fractured everywhere, every country presenting its own complex, deeply embedded hurdles.

AI brought the allure of scale, automation and intelligence. Investors saw a billion-dollar market spanning borders and they were ready. But in a system this complex, this entangled, only something truly transformative could make a meaningful dent.

The antidote doesn’t begin with a cure, but rather with a deep-seated exposure. Before we could imagine fixing the system, AI forced us to see just how fractured it really was.

Early days of building with AI

In the early days of building with AI, it didn’t feel like working with some slick, cutting-edge machine. It felt like stitching together a glitch-prone, barely-coordinated organism with duct tape, intuition, and sheer persistence.

Y Combinator (one of the world’s most prestigious startup accelerators) invited us to San Francisco, back when Sam Altman still served as president and Paul Graham’s rallying cry to be “Relentlessly Resourceful” adorned every startup’s wall. We were encouraged to roll out our solution within developed economies, but we were the dewy-eyed, idealistic founders — determined and focused on emerging economies. All the while avoiding overly-competitive or saturated markets. A strategic move, we thought.

India took the trophy for the most complex problem statements by far. 

In the early days, I remember the quote one of our clients shared, he said: 

“India has mouthwatering opportunities and eye-watering challenges”.

The patient to doctor ratios were abysmal. We found that families were going bankrupt (emotionally and financially) just trying to get care: myriad of health issues, out-of-pocket health expenses, convoluted claims processes, insurance systems that felt more like traps than safety nets. 

But here’s where the naiveté showed: knowing where the system was broken doesn’t mean you can fix it. Not easily. AI could diagnose failure points with elegance, even offer a path to optimisation. Implementation, though, was messy! Especially in large, legacy-bound infrastructures. It became less about solving real-world challenges through tech and more about navigating change at the organisational, structural, and human level.

We saw the interconnectedness of how easily one departments fix can become another's hurdle.

In other words, growth in one area highlighted inertia in another:A great idea for onboarding customers led to decades-old underwriting guidelines being challenged. A triage system that prioritised patients walking into the hospitals, yet the doctor’s time became a bottleneck.

Imagine, a real life whack-a-mole. 

Note: The following sections seek your full presence. We’re about to explore the depth and complexity of building an AI system from the ground up, explained in (relatively) straightforward terms. Along the way, we’ll uncover some of the less visible layers behind the AI systems shaping our world today. I invite you to read not just the words, but the spaces between them.

From Scrappy to Significant

We walked into the era of building AI systems in geographies where data wasn’t handed to us on a silver platter to build models with — no clinical records, no structured inputs, no ideal training sets to plug into. Just fractured processes, clinical notes buried in hand-written sludge, and the urgent need to make sense of it all. 

It was a modern-day mechanical turk. Every step manually validated by a human clinician behind the scenes. We started scrappy, almost artisanal. 

We used simple but surprisingly powerful methods. Imagine teaching a machine to read stacks of messy notes and pull out the most important words, then “predict” or “guess” which category they belonged to — like a child sorting coloured hats into the right sections, it began sorting medical information where 'breathing difficulties' was tagged for respiratory issues or ‘chest pain’ to the cardiovascular system. It was taught to notice when certain events tended to follow others, like fever and headache leading to cough. These early methods were rough approximations of how a junior doctor reasoned through symptoms to diagnosis.1

As the models and their complexity grew, they learned to hold context. Instead of looking at each symptom in isolation, they could connect the dots across time — like a seasoned doctor piecing together a patient’s history. From there, the models evolved into ones that could grasp relationships.2

On a broader scale, recognising that a certain drug is linked to a certain disease, or certain combinations of conditions increase the chance of co-morbidities in the future. 

The system learned by reading enormous amounts of medical text, everything from handwritten doctor notes to research papers and medical books. Using Natural Language Processing (NLP), it could pick out key details and connect them.

For example, just by reading, it could now recognise that Metformin is used to treat Type II Diabetes. It would tag “Type II Diabetes” as a condition, “Metformin” as a medication, and mark the connection between them as “strong” because it noticed the repetitions in the data. This is a simple example, but in its representation it shows a huge leap in teaching a machine to understand medical knowledge in context. Something that may feel archaic in comparison to the current capabilities of AI systems now.

An Interconnected Clinical Brain

By 2018, this leap made it possible to create what we called a Medical Knowledge Graph: a vast digital map of how diseases, symptoms, treatments, and medications relate. 

One of its kind for emerging economies!

Think of it as a living, interconnected brain constantly updating itself as it learned. 

The progression from naive statistical tools to contextual learning models marked a turning point — mapping the cognitive intelligence of specialist clinicians into one vast system.

The impact was clear: in places where reliable data barely existed, we could now generate meaningful insights. What began as crude text sorting grew into context-aware intelligence that supported doctors, insurers, and policymakers — it could now predict individual health conditions, help doctors make more precise diagnostic and treatment decisions, support insurers with tasks like onboarding customers, underwriting risk, and even refining actuarial calculations. In short, an evolved system that could support the entire lifecycle of complex industries.3

Advancing AI Models with Limited Data

You’re probably thinking, if there wasn’t enough data to train the models, how did we do it?

Enter → Synthetic Datasets

Even with the Medical Knowledge Graph, the system was hungry for more data. You can think of it like a highly intelligent student who had mastered the textbooks but needed real-world cases to keep improving. 

So we built synthetic datasets. 

PS: Synthetic ≠ Fake: These are realistic, evidence-driven datasets that reflect the complexity of real-world scenarios. Synthetic simulations are not new age — they’ve existed since the mid 1900s.

Monte Carlo methods, developed in the 1940s during the Manhattan Project, were among the earliest uses of synthetic data (probabilistically generated samples) for simulations in nuclear physics. These methods used random sampling to simulate complex physical phenomena where real-world experimentation was impossible or too dangerous. Monte Carlo methods are now core to simulations, probabilistic algorithms, and synthetic data generation.

It meant we could now train systems where data scarcity once halted progress.

And, while these weren’t real patient records, they were designed meticulously to look and behave like the real thing. Instead of pulling from scarce or private health files, we essentially taught the AI to invent realistic scenarios: a patient with a combination of symptoms, lab results, and medical history that a real doctor might encounter.

Imagine feeding the system a prompt like, “A 45-year-old patient with diabetes presents with shortness of breath.” The AI would then generate a plausible set of test results, risk factors, and treatment paths based on everything it had learned from the Medical Knowledge Graph to emulate hundreds of thousands of cases. In this way, it could create both individual patient profiles and digital populations (calibrated on real geographical data based on disease incidence & prevalence).4

The advantage was huge. We could now train sophisticated models even when partner hospitals or insurers didn’t have clean or complete data. These synthetic cases gave us a cold-start strategy: the models could perform well from day one, without waiting years to accumulate real patient data. And with real-time feedback loops, the system kept refining itself just like a doctor improving with experience.5

Now, you’re probably wondering how is this accurate and how well does it actually reflect the real world?

  1. Medical Knowledge Graph as the Foundation

  • These synthetic patients aren’t created out of thin air. They’re anchored in a Medical Knowledge Graph, which encodes verified clinical relationships:

    • Symptoms ↔ Diseases

    • Risk factors ↔ Outcomes

    • Medications ↔ Side effects & contraindications

  • This ensures that even though the “patient” is generated (i.e. simulated), the logic of their condition mirrors real-world clinical pathways.

    2. Incidence & Prevalence Anchoring

  • The digital populations are not random. They’re calibrated against epidemiological data for a given geography (e.g., prevalence of diabetes in rural India vs. urban India varies).

  • This gives the generated population statistical realism, so the distribution of conditions, complications, and risk factors reflects what clinicians actually see.

    3. Cold-Start Validation via Clinician Feedback

  • Doctors and medical reviewers are part of the feedback loop. (You’ve probably overhead the term RLHF - Real Life Human Feedback, this is what it is. Having a human in the loop to validate the outputs of the models).

  • When the AI generates a case (say, a diabetic patient with a kidney disease), clinicians can flag whether the lab values, medication patterns, or risk profiles make sense.

  • These feedback loops function like residency training for the AI: it keeps adjusting until its scenarios are clinically sound.

    4. Dynamic Updating with Real Data

  • As real patient data becomes available, the system uses it to fine-tune and validate the synthetic cases.

  • Over time, the synthetic data serves as scaffolding allowing the model to start strong while real-world feedback strengthens accuracy.

Let’s elaborate with an analogy: think of it like a flight simulator for pilots.

1950s onward, synthetic simulation became foundational in aerospace engineering (e.g., flight simulators, trajectory prediction) and military strategy (e.g., radar systems, weapons testing). These used generated data to mimic real-world conditions under controlled variables.

  • The simulator doesn’t use a “real plane,” but it’s built on the exact physics, weather models, and cockpit controls of real aviation.

  • A pilot trained in the simulator emerges competent and ready for real skies.

  • Similarly, synthetic cases train the AI in realistic conditions, ensuring safe and effective decision-making when it encounters actual patient data.

So, was it perfect? No. But it was revolutionary! Synthetic datasets became the foundation that allowed us to break through the barrier of data scarcity and build clinical models where others had stalled.

Emergent Behaviours in AI

The real breakthrough didn’t come from a single “eureka” moment, but from what the system started doing once all the pieces worked in tandem.

By 2019, we began to witness emergent behaviour. Not just better outputs but insights. The models started surfacing patterns we didn’t label, spotting edge cases, drawing parallels that weren’t in the training set. 

They began reflecting associations more aligned with practitioner intuition than with traditional model outputs.

With advancement in unsupervised learning (where the AI learns patterns on its own without specific labels, finding its own structure within the chaos) and deep generative techniques, we were closely observing this emergence — models learned patterns, extrapolated meaning, identified contradictions, made connections that no linear system ever could. Logic began surfacing from abstract layers. Not programmed. Just... found.

For example, our clinical team was stunned when the system identified disease characteristics it hadn’t been explicitly trained on (something a doctor would have missed without invasive tests), or when it suggested preventative health plans that reduced avoidable insurance claims (predicting future high-cost events with surprising accuracy). It wasn’t magic — the system was learning from patterns buried deep in the data, patterns no human could have pieced together at that scale.

It was exhilarating. And unnerving. Because while the predictions were increasingly accurate, the how became opaque. We’d crossed the threshold into systems that gave us insights we couldn’t always explain. 

This is where the ‘aha!’ moments started to unravel. 

And, what started as models that detected diseases soon began to mutate, adapting itself to new domains entirely. Expanding its skillset on its own.

From diagnosis to diseasemanagement, figuring out optimal treatment paths.
From health-risk prediction to nation-wide health indicators.From spotting fraudulent claims to avant-garde policy design.

And, so we built… 

Training models with limited data.
Continuing to teach machines to apply the cognitive processes of a clinician, an underwriter, an actuary, a sales personnel — you get the point.

Slowly, we began building capacity into these systems, teaching the models to need less and less input. 

Explainability and Black-Box Models

Mystery might inspire wonder, but in healthcare, it’s malpractice.

As powerful as these systems became, one of the biggest challenges was explainability. It wasn’t enough for the system to say what it concluded; we needed to understand how it got there. Especially within healthcare & financing, industries riddled with regulatory burdens, where accountability cannot be deferred to a black box. 

Explainability became non-negotiable. 

Some of the most unexpected breakthroughs arrived not from the outcomes the models produced but from the meta-models we built to decipher those outcomes. 
Think of it like models designed to explain models.

They could show which pieces of data influenced a decision, outline the reasoning path they took, and translate that into plain language. 

For instance, which symptom carried the most weight, which lab results were critical, the historical data that was mostrelevant… Instead of just outputting a prediction, the system could walk doctors, insurers, and regulators through its logic step by step — bridging the gap between complex computational systems and human understanding.

We became known for the explainability behind our models. While our competitors were still building black-box, stealthy intelligence. 

At first, this felt like a victory. Clients loved the transparency. Boards signed off with more confidence. But the deeper we traced those paths, the more uncomfortable truths surfaced. 

The models were explaining medical decisions, but they were also lighting up the fault lines, exposing how fragmented and misaligned the ecosystem was. 

Departments working in silos. Incentives for doctors, hospitals, insurers were pulling in opposite directions. Sometimes their goals were actively working against each other and against the patient’s (or customer’s) best interest.

We called it an ontological fracture — a fundamental break in how the system should work versus how it did work. 

For example, imagine patient treatments being prolonged to maximise bed occupancy at hospitals or insurance sales teams were subtly asking customers to withhold sharing medical information to bypass underwriting rules.

We became the accidentally auditors, the unsolicited consultants no one had asked for. 

There were many awkward meetings. Suddenly executives were grappling with a map of dysfunction they couldn’t ignore. The buzzing rooms grew silent and the tone of the questions shifted. 

And it became clear, explainability was challenging organisations to confront their own deep-seated, systemic fractures.

The Vantage Point

So, why am I telling you all this?

This experience provides a pragmatic vantage point on AI that differs sharply from the mainstream narrative. While the world is now inundated with AI tools that produce fluent outputs and novel applications, but as someone who has lived through the emergent AI era, I’m inspired to share a more grounded and critical understanding.

Moving Beyond Outputs

Most people focus on what AI produces: a convincing essay, a medical summary, or a customer service reply. I want you to think about how the system arrives at those outputs. 

Reasoning ≠ Explainability(we’ll dive deeper into this in part 2 — soon to come)

Watching models generate insights not explicitly trained into them forces a shift from surface-level fascination to a deeper inquiry into the mechanics of intelligence. We no longer ask only, “What is the model output?” but, “How and why is the model outlining this pattern?”

Early Exposure to Opaque Intelligence

When AI is making decisions about people’s health and their finances, it forces you to encounter the challenge of accuracy without transparency. This creates a lived understanding of what many are only beginning to grapple with: explainability is as important as accuracy. It’s a point of tension — we want the benefits of precision, but we also need frameworks to interpret and trust these insights.

Grounded Awareness of Risk and Responsibility

Many of us currently encounter AI as a novel productivity tool, I’ve seen its decisions touch lives in many high-stakes contexts. That builds a sense of responsibility, innovation must be balanced with safeguards. I know from experience that unchecked enthusiasm without governance leads to risk — not just technical risk, but human and institutional risk.

Moving Forward

The broader world may be dazzled by AI’s fluency, speed and charm. But, I invite you to cut into something deeper: we’ve already seen the limits of human oversight when AI begins surfacing patterns beyond explicit training. Instead of being swept up in hype cycles, we need a more sober, strategic view. Can we ask the harder questions: How do we validate what we can’t fully explain? How do we govern what we can’t fully predict?

  • Building frameworks of trust: advocating for explainability, transparency, and audit mechanisms in AI deployment.

  • Educating stakeholders: helping leaders understand that AI’s outputs are the tip of a much deeper process that will magnify and reveal the existing cracks.

  • Balancing innovation with accountability: showing that rapid deployment must be matched with responsible guardrails.

  • Shaping public discourse: grounding conversations in lived realities, practical applications, not just speculative hype or doom & gloom fear mongering.

The bottom line is this:
If we deploy AI into systems already buckling under the weight of a massive global transition, it will do one of two things: amplify existing fractures or relentlessly pursue the outcomes it was trained for — regardless of consequences.

But if we’re willing to face the uncomfortable truths and position AI as a tool to expose what’s no longer working, we give ourselves a chance to meet what’s coming not with patchwork solutions, but deeply rooted, systemic resilience.

I hope this perspective equips you to navigate today’s AI influx with pragmatism, foresight, and responsibility. Don’t just be amused by the surface while missing the depth of what’s unfolding.

Part One was about discovering the scale of the problem and the unexpected ways AI exposed it. Part Two will go deeper: what happens when those revelations collide with entrenched power, regulation, and the human fear of change?

Until next time.

References & Notes:

1

  • TF-IDF (Term Frequency - Inverse Document Frequency) to rank relevance & find keywords — a statistical shot in the dark. 

  • Naive Bayes Models look at words in a document and makes a guess about what category it belongs to, for e.g.: “breathing difficulties” would link to respiratory over cardiac. We used this for crude categorisation — limited, but something. 

2

Markov chains & RNNs (Recurrent Neural Networks) marked sequential patterns for things happening in a certain order. RNNs possess a “memory” function where past information can be used to predict the next contextual sequence (e.g.: Is fever followed by cough?) — brittle, relational, but serviceable.

3

This continuous data enrichment and scaling of the graph result in complex information networks used for various client use-cases. The Neural Networks and Deep Learning methodologies were built on top of the graph. Some of the algorithms that made this possible:

  • Biomedical Natural Language Processing (NLP) algorithms These were used to extract and mine key phrases from medical discharge notes and to enrich datasets for modelling.

  • Neighbourhood Preserving Graph Embeddings: Fancy name but it basically helped to fill the gaps within the medical knowledge graph by predicting noisy and missing links, identify anomalies and improve data quality in data scarce geographies.

  • Ontology-guided Data Augmentation: These are structured dictionaries for medical concepts. Deep learning algorithms combined with medical ontologies were extracting key phrases for conditions, investigations, treatments, etc. and meeting regulatory needs. 

  • Contextual Clinical BERT Embeddings: Clinical BERT models were used to detect various medical entity types in plain text, including diseases, symptoms, risk factors, lab values, investigations, procedures, medical events, and their characteristics.

  • ML-driven Context-Relevant Dialogue System: This ChatGPT-like system was developed to mimic a doctor's symptom history-taking process and predict a diagnosis, which is now part of the Quro product used in multiple client chatbot use-cases.

4

Graph Search-Based Patient Case Simulation Algorithm: A novel algorithm was developed that generates realistic digital cases directly from the proprietary medical knowledge-base. Now, that cognitive intelligence of a multi-specialist clinician was put to good use. The brain we built was used to invent plausible patient data. This algorithm led to significant research contribution in IEEE, AAAI, ACM, etc.

5

  • Continuous Improvement with Feedback Loops: The system incorporates a real-time closed-loop feedback system (think: RLHF i.e. Real Life Human Feedback) to continuously evolve large-scale clinical knowledge bases and the lifecycle of health prediction models. This indicates an iterative and adaptive approach to algorithmic refinement.

  • Mimicking Real Patient Records and Overall Population: The synthetic simulation workbench is designed to simulate data that accurately mocks both real patient records and overall population trends. This comprehensive simulation capability allows for micro-level and macro-level healthcare profile generation.

  • Addressing Data Scarcity and Imbalance (Cold-Start Approach): We highlighted a "novel cold-start (built from the ground up) approach" that mitigates the risk of poor quality or insufficient data. This meant that the pre-built and validated models have no dependency on partner-specific data, enabling faster go-to-market for partners.

  • Advancing Evaluation Methodologies: The developed algorithms also advance the field of evaluation methodologies and propose novel metrics for evolving clinical knowledge bases and health risk prediction models. This commitment to robust evaluation set our approach apart

Previous
Previous

The Dialectic of AI and Authorship