Architecture review
One week inside your system. Retrieval strategy, evaluation design, cost and latency, failure modes. Ends in a written recommendation and a plan — not a deck.
Retrieval systems, agents and ML platforms — architected, built and shipped. Most AI projects die between the demo and the deploy. That gap is the entire job.
Most engagements start small. A review that becomes a build is better for both sides than a build that never gets signed.
One week inside your system. Retrieval strategy, evaluation design, cost and latency, failure modes. Ends in a written recommendation and a plan — not a deck.
Fixed scope, shipped and deployed: a retrieval system, an agentic application, an LLM feature, an ML pipeline. Evaluation harness included, and handover documentation your engineers can actually work from.
Agents and workflows that run part of the business, built on your stack, with a human approval step wherever it touches a customer.
We act as your Head of AI. Architecture calls, model and vendor choices, evaluation and governance, code review, hiring input. Hands-on where it matters.
The evaluation harness exists before the demo does. Click any stage — this is the actual shape of a retrieval system we'd ship.
Human review sits on every path where a wrong answer has a real cost.
Client names appear only where we're free to use them. Every number here is measured, not estimated.
Cloud-native ML pipelines on AWS serving real-time medical assessments at 99.5% uptime. A deep-learning wound-classification system reaching 82% accuracy in clinical testing, with the evaluation and governance practice built around it — including human-in-the-loop clinical review, so a clinician could overrule the model and we could see when they did.
An agentic conversational system on LangGraph with Elasticsearch-backed retrieval, doing autonomous multi-step reasoning and dynamic tool orchestration across government services. Alongside it, a data-warehousing platform for Qatar's Ministry of Commerce and Industry integrating large-scale trade data for automated regulatory compliance reporting, and predictive modelling with scenario analysis for Kuwait Fund budget forecasting.
A resume–job matching system on a Vespa vector database processing 1,000+ pairs daily for 30K+ monthly active users from underserved communities, plus the scraping pipeline feeding it — 500+ postings a day with transformer-based parsing and skill extraction.
An audience-targeting platform over 20M+ ad IDs and 1.8B+ captured signals, with distributed PySpark pipelines on Databricks tuned for that volume. Elsewhere: demand forecasting and inventory optimisation for Tory Burch's international shipping, an enterprise data warehouse with Tumi, and a churn model with an automated weekly pipeline that cut attrition by 18%.
Three tracks, from executive strategy to advanced engineering. Delivered in Arabic or English. The first two sessions of every track are free — you commit only if it's good.
Find the high-yield use cases, calculate real ROI, govern your data, and lead AI engineers without a technical background.
Zero-to-one on modern AI pipelines: embeddings, vector databases, custom RAG systems, and live deployment.
Multi-agent architectures, fine-tuning open-weight models, clinical-grade computer vision, and cost and latency at enterprise scale.
A real engineer reads every one of these, usually the same day. The more specific you are about the problem, the more useful the first reply will be.
Rather talk? Book 30 minutes — no pitch, and if we're not the right fit we'll say so.