Skip to content
Currently available — for the right work·France903+ Day French Streak·2026 Q2 calendar — open now
Microsoft AI · Applied AI Lead, Health · London

Make frontier modelsuseful, trusted, and safeacross people's health journeys.

Microsoft AI's Health team is hiring an Applied AI Lead to turn frontier models into Copilot Health — building rigorous health evaluations, architecting agentic LLM orchestration, and growing a team of Applied AI Engineers while staying hands-on. I bring a decade of high-agency, player-coach delivery, production agentic systems with eval harnesses (Claude-led, Gemini second), and a HIPAA clinical platform behind me — and London is where my family already is.

Ask Fauzul
Eval-gated delivery
9+ yrs
hands-on founder & player-coach leadership
ELO — co-founder → CIO
3,242
patients on the HIPAA clinical platform I architected
EKAGRA — HIPAA EHR
4×
faster time-to-market via CI/CD & quality discipline
measured at ELO
How the work maps

Eval-gated delivery. Agentic orchestration. Health at the core.

Agentic run · NewScriber
Intent
Retrieve
Draft
Eval gate
Ship
Accuracyevery claim checked against source before synthesis
Safetygated behind review before anything ships
Utilityevals decide what ships — not vibes
Ship only on pass — a failed gate loops back to Draft.
Illustrative — the gate every claim on this page runs through. No scores shown because none are invented.
Trusted

Health evaluations that drive product decisions

I build eval pipelines where 'feels right' is never the release criterion — grounding evals gate 's scripts before synthesis, and grounds every recommendation against live sources. Paired with a HIPAA clinical platform at , that is the accuracy-safety-utility bar this seat needs, held on real receipts.

RetrieveReasonVerifyParseSynthesize
Illustrative — the hub-and-spoke shape behind VisaPros' six-agent A2A architecture.
Orchestrated

Agentic multi-step systems, built to hold in production

Harness and context engineering, tool use, retrieval and agentic multi-step orchestration are daily work — (six-agent A2A on Google ADK), (multi-agent ReAct on n8n), and the Claude CMS behind CI/CD gates at . Model families are blended per task, never bolted on.

Compounding

Engineers grown into leaders, not just features shipped

Weekly design and code review, 1:1 coaching, judgment over throughput — alumni I mentored now lead engineering in Norway (2) and at two top local firms, with others coached into graduate study in Canada and a UX design career in Germany. Leading and growing a team of Applied AI Engineers while staying a credible authority on evals is the exact loop I run.

EKAGRA — HIPAA clinical platform

In production. Under HIPAA. Not a demo.

64doctors on the platform
3,242patients
13,340schedules
7centers
HIPAAGDPRWCAG 2
The signature discipline

Together we make frontier models useful, trusted and safe across people's health journeys — with evaluation and orchestration rigorous enough to earn that trust.

I build the systems that turn frontier models into products people can trust — then I keep raising the bar on evaluation, orchestration, and how the team ships.

Fauzul — on why Microsoft AI, Health

Responsibilities 11
StrongTurn frontier models into products people trust with their health — build rigorous, health-specific evaluations that drive real product decisionsThis fuses the two things I already do: build eval-driven AI systems and ship in health. NewScriber's grounding evals gate every script before synthesis; VisaPros grounds each recommendation against live sources; and I architected a HIPAA clinical…
StrongMaster orchestration — harness and context engineering, blending model classes and families, applying state-of-the-art techniques — and bring bleeding-edge AI expertise that uplevels the teamI already blend model classes and families per task — Claude (led), Gemini, Kimi via OpenRouter, on-device…
StrongHands-on technical leadership: predominantly build Copilot Health, bridging the latest research and product to establish Copilot as a leader in safe, trustworthy health informationBridging frontier technique and shipped product under a trust bar is my default seat: agentic capstones and hackathon…
StrongLead, mentor, and grow a team of Applied AI EngineersWeekly 1:1s and review culture as default across a distributed team. Alumni now lead engineering in Norway (2) and at two top local tech firms; others I coached stepped into graduate study in Canada and a UX design career in Germany — mentoring as a…
StrongStay deeply hands-on — set the bar through code and design reviews, lead on the hardest problems, and remain a credible authority on evaluations and LLM systemsCo-founder → EM → CIO who still ships daily in Python, TypeScript and Go — roadmap influence and code review in…
StrongCo-own the roadmap with product leads — qualify and size opportunities, co-author the roadmap, and lead architecture and 0→1 developmentDecade of roadmap ownership under founder titles — qualifying fuzzy briefs, sizing them, and leading architecture…
StrongOwn delivery — plan and prioritise the roadmap, bias to shipping and learning against a high-quality bar, and ensure production reliabilityBias-to-ship with a quality bar is the operating system: 4× faster time-to-market (75% cycle-time reduction) via CI/CD and repeatable workflows, agent-authored changes gated by verification, and production reliability owned through support and escalation.
StrongDefine evaluation strategy — systems that test LLM capabilities in health, including internal benchmarking and regression testing for accuracy, safety, and utility; interpret and communicate results to stakeholdersEvaluation as a first-class surface is how I decide what ships — grounding-eval gates in NewScriber and VisaPros, plus eval/harness hill-climbing as a named skill. I communicate results plainly to non-technical stakeholders (a decade of client-facing…
StrongArchitect LLM orchestration — agentic, multi-step systems combining prompt/context engineering, tool use, and retrieval; champion best practices for reliable deployment at scaleDeep hands-on: VisaPros is a six-agent A2A hub-and-spoke on Google ADK (sequential parse → parallel specialists →…
StrongRun and direct experiments on prompting and orchestration techniques against internal and industry benchmarks, turning findings into product improvementsIterating prompts, harnesses and orchestration against measurable signals is habit — hill-climbing until a system…
AdaptableInvest in tooling — improve internal evaluation tooling and data pipelines, including dataset sourcing, curation, and synthesisStrong on eval tooling and ingestion pipelines — grounding harnesses, Firecrawl datascape ingestion, n8n workflows, a Sharp image pipeline over 200+ sources at . Text-dataset sourcing, curation and synthesis at scale for evaluation is…Adapt: Own the eval-tooling and harness layer immediately; ramp on large-scale dataset curation and synthesis for health evals by pairing with the team's…

Health-specific evals.Agentic orchestration.A team that ships safely.

Required qualifications 6
AdaptableBachelor's or higher degree in Computer Science or a related technical discipline, AND significant Python programming experience / ML researchPython is strong and applied — agentic systems, FastAPI services and data work in production, backed by Kaggle's Python for Data Science — on top of a decade of shipped applied-AI engineering. Field of Study: **Computer Science & Engineering · North South…Adapt: The operative strengths — a decade of shipped applied-AI work, strong applied Python, and 8+ years leadership — are present. Research-adjacent…
StrongVery strong proficiency designing, building, and running LLM evaluations — eval pipelines, dataset curation/synthesis, automated analyses, and explaining results to stakeholdersEval pipelines that gate production are core: NewScriber grounding-checks every script against source bodies…
StrongVery strong proficiency with LLM orchestration — deep hands-on building with and around LLMs: prompt/context engineering, tool use, harness engineering, retrieval, agentic multi-step systems, and tools to analyse performanceThis is the centre of my recent work: VisaPros (A2A multi-agent on Google ADK), NewScriber (ReAct + n8n…
StrongProven engineering leadership — 8+ years of software engineering including 3+ years leading technical teams or projects as a tech lead and/or people manager9+ years since co-founding , managing people for 7 of them while staying hands-on — well past the JD's 3-year bar — formally Software Engineering Manager (Dec 2022 – Sept 2024) then CIO (Sept 2024 – present) — leading a **22-person…
Strong0-to-1 experience with a bias to shipping and learning while balancing a high-quality bar0→1 across five verticals — music-tech, gaming, hospitality, health-tech, professional networks — with weekly cadence over quarterly perfection and a validate-before-overbuilding instinct (four Haiba pop-ups before committing capital; ** shipped…
StrongExperience collaborating in cross-functional teams, working through ambiguity, and contributing to a positive, inclusive environmentCross-functional by construction — product, design, engineering, QA and business in the same room at and in…
No claim without a receipt

Every verdict above maps to shipped work. Pick one and challenge it.

Preferred qualifications & location 7
StrongHealthcare technology or health-domain experienceEKAGRA is a genuine health-tech engagement — a HIPAA clinical platform for wound care, diabetes and nephrology where I architected the patient-data pipeline and EHR integration (64 doctors, 3,242 patients, 13,340 schedules, 7 centers), aligned to…
AdaptableData engineering — sourcing, curating, and processing text datasets at scaleStrong on data-intensive backends and ingestion pipelines — PostgreSQL, Firecrawl datascape ingestion, n8n…
StrongTranslating cutting-edge research into shipped products in a fast-paced, startup-like environmentThat translation is the through-line: a Google × Kaggle capstone became VisaPros; a one-week Milan AI Week sprint…
StrongPassion for conversational AI and its deploymentNewScriber's dual-host conversational audio (multi-speaker turn-taking, interruption, continuity), TagRamp's agentic warm intros, and this site's on-device conversational chat (Gemma 4 E2B on WebGPU) — conversational AI is something I build and ship…
StrongStrong written and verbal communication with cross-functional teams (PMs, designers, engineers)UX-engineer origins plus founder/CIO seat — I speak product, design and engineering without a translation layer…
StrongPassion for learning new technology and staying current on AI trends, best practices, and emerging patternsFrontier-by-default: agentic hackathons, first-party Anthropic and Google × Kaggle certifications, and a…
StrongLondon, United Kingdom — 4 days/week in-office (travel <25%)London is a deliberate choice. The densest frontier-tech hub outside the US West Coast — AI labs, the NHS and health-system programmes, financial services and scale-ups within one commute — which is exactly the environment for applied-AI work in health…

Top strengths

What lands strong, on receipts.
Evaluation + orchestration discipline in production
Grounding-eval-gated pipelines (), an A2A agentic mesh (), and a Claude CMS behind CI/CD gates () — the JD's two 'very strong proficiency' bars, already how I ship. Anthropic Claude Code 101 + Introduction to Agent Skills and the Google × Kaggle AI Agents Intensive as dated proof.
A genuine health-domain edge
— a HIPAA clinical platform (64 doctors, 3,242 patients, 13,340 schedules, 7 centers) where I architected the patient-data pipeline and EHR. Health-tech depth is rare for an applied-AI lead; here it is real.
Player-coach leadership, still hands-on
Co-founder → EM → CIO who still ships Python, TypeScript and Go — roadmap and code review in the same week. Alumni now lead engineering in Norway (2) and at two top local firms; others grew into graduate study in Canada and a UX design career in Germany.
Model-agnostic by practice, honest by default
Blends Claude, Gemini, Kimi via OpenRouter and on-device Gemma per task; the right tool wins the task, and Claude leads production client delivery today. The discipline transfers to whatever Copilot Health runs on.

Adaptable ramps

Named, not buried.
'Significant ML research' as a researcher-track credentialOngoing
Applied-AI depth is real — production LLM systems, eval harnesses, agentic orchestration. Would own the applied build and pair with research-track colleagues on formal health benchmark design. Grounded by the Google × Kaggle AI Agents Intensive and Microsoft Azure ML AutoML training — named honestly, never overclaimed as a publication record.
Text-dataset sourcing, curation, and synthesis at scale (data engineering)First quarter
Data-intensive backends, ingestion pipelines (Firecrawl, n8n) and eval loops transfer directly; deepen large-scale dataset curation and synthesis for health evals with the team's data-engineering practices.
Completed CS degreeN/A — experience-based
Field of Study: Computer Science & Engineering · North South University (2012–2015). Qualification is experience-based — a decade of shipped applied-AI and leadership work, which the 8+ years leadership requirement speaks to.
Zero hard gaps
Every requirement is strong or a named, bridgeable ramp.
The objections, answered straight

Questions a sharp hiring loop would raise.

The objections a sharp Applied AI Lead loop would raise — including the AI-stack one — answered straight.

Copilot Health runs largely on OpenAI and other models. Your production stack is Claude-led — is that a problem?

The value is the discipline, not the vendor.

No, and I won't paper over it. My hands-on production AI has been Claude-led with Gemini second — I pick model and harness per task, not by vendor loyalty, and haven't had reason to run OpenAI's models in production yet. But the value here is the discipline, not the vendor: harness and context engineering, grounding evals, tool use, retrieval, agentic multi-step systems, and benchmarking are model-agnostic and transfer to any family. I already blend model classes and families per task — Claude, Gemini, Kimi via OpenRouter, on-device Gemma — so getting productive on Copilot Health's stack is a ramp measured in weeks, not a re-education.

The role asks for significant ML research. You're an applied-AI engineer, not a researcher — right?

Correct, and I won't pretend otherwise. I ship applied LLM systems and evaluation harnesses; I have not published ML research. What I do bring is research-adjacent grounding — the Google × Kaggle 5-Day AI Agents Intensive and Microsoft's Azure ML AutoML curriculum — plus the ability to turn cutting-edge technique into shipped product fast. On this team I'd own the applied build and health-eval design, and partner with research-track colleagues where formal research is the right tool.

Your leadership is founder/CIO at a studio, not a big-tech TLM. Can you lead an Applied AI Engineer team at Microsoft's scale?

I have 8+ years of software engineering and have managed people for 7 of them while staying hands-on — formally Engineering Manager 2022–2024, then CIO — leading a 22-person cross-functional core and dedicated client pods, including hiring and HR for a Norway team. Alumni I mentored now lead engineering in Norway (2) and at two top local firms, with others coached into graduate study in Canada and a UX design career in Germany. The JD explicitly says a strong tech lead ready to step into a TLM role will be considered; I am past that bar, and I stay hands-on rather than drifting into pure management.

Have you built health-specific LLM evaluations before?

I've built the two halves this role fuses: a HIPAA clinical platform ( — patient-data pipeline and EHR) and grounding-eval pipelines that gate production (, ). Health-specific benchmark suites for accuracy, safety and utility are what I would build first — on real eval discipline and real health-domain context, not from a standing start. That combination is exactly why this seat fits.

Do you have a completed CS degree?

Field of Study: Computer Science & Engineering at North South University (2012–2015). I do not claim a completed degree. My qualification is experience-based — a decade of shipped applied-AI systems and 8+ years of engineering leadership are the operative credentials, and strong applied Python and LLM engineering back the technical bar.

This is 4 days a week in-office in London. Are you genuinely relocating?

Yes — and it isn't a cold bet. My sister lives in London, so it's a real personal base, and my professional history in Europe is genuine: multi-year London B2B and audio-engineering clients, a 6+ year Zürich engagement (, acquired by Loopcloud 2026), and a Norway fintech pod (Frontgo, Vipps). Four days a week in-office is exactly where I want to be for this seat.

The stack behind the claims

Claude-led, model picked per task, not by vendor loyalty.

The stack and receipts behind the claims — Claude-led, Gemini second, model picked per task; eval harnesses and agentic orchestration as habit.

Evaluation & orchestration
LLM eval pipelines & grounding gatesHarness & context engineeringAgentic multi-step (A2A, MCP)Tool use & retrieval (RAG)Prompt engineeringToken & cost-aware context engineering
AI & models
Claude Code (Anthropic-certified)Claude Skills, hooks & subagentsGeminiGoogle ADKKimi via OpenRouterOn-device Gemma 4 E2Bn8n orchestration
Languages & platform
PythonTypeScriptReactGoNode.jsPostgreSQLDockerAzure (Gemini TTS)Google Cloud RunCI/CD
Health & compliance
HIPAA EHR (EKAGRA)Clinical data pipelinesGDPRWCAG 2Compliance as a product surface
Leadership & delivery
Player-coach engineering leadershipTeam mentoring (Norway ×2, Canada, Germany)0→1 deliveryCSPO / roadmap ownershipCross-functional stakeholder alignment
Applied AI Lead, Health · Microsoft AI

let's make Copilot Health_safe enough to trust.

London is where I want to build, and that is a professional judgment before it is a personal one. It is the densest frontier-tech hub outside the US West Coast — the place where the AI labs, the banks, the NHS and public-sector programmes, the scale-ups and the enterprise buyers all sit inside the same hour. For an engineer who works *embedded with customers*, that concentration is the whole point: more real businesses to talk to, more industries in one commute, and the fastest way to keep learning how technology and business actually move each other. The energy is real too — you feel the pace of the city in the work. It is also a genuine personal base: my sister lives there, and my delivery history in Europe is real — multi-year London client work, a 6+ year Zürich engagement (Jamahook, acquired by Loopcloud 2026), and a Norway fintech pod (Frontgo, Vipps). See London for the full picture. This seat is 4 days a week in-office and I want that — being in the room, in that city, is the point.

Application fit for Member of Technical Staff — Applied AI Lead, Health at MicrosoftMicrosoft hub

A personal application fit page, not affiliated with or endorsed by Microsoft. Copilot / Azure / Fabric / ISD references are illustrative of the role's own team and domain — no confidential or internal Microsoft information is shown.