Shaibi Shamsudeen

Open to senior AI architecture roles.

AI systems that connect models, data and controls to real work.

I'm Shaibi Shamsudeen, a Systems Architect and AI safety researcher. I've shipped multi-agent research tools, conversational platforms for real-estate brokerages and ML trading pipelines, and led the functional architecture of a risk platform later acquired by Dow Jones.

The stack I ship with

  • Python
  • LangGraph
  • LangChain
  • ChromaDB
  • OpenAI
  • Gemini
  • Hugging Face
  • PyTorch
  • TensorFlow
  • FastAPI
  • React
  • TypeScript
  • PostgreSQL
  • Supabase
  • AWS
  • Docker
2nd
Place at the Outskill AI Accelerator, from 250+ participantsLuminar
<1%
Prediction error after the correction engine I designedTrading systems
40%
Efficiency gain in compliance workflowsGRC automation
2
Accepted LLM safety workshop papers, NeurIPS 2025 and COLM 2026Both papers

// Selected work

Systems I've designed and shipped

Each project links a real problem to a system design, my specific contribution and an outcome you can check. Filter by the layer of an AI system you care about.

Open source

Multi-agent research system

Luminar

45% better retrieval efficiency than the project baseline. Second prize among 250+ accelerator participants.

A deep-research system that coordinates specialised agents across web, video, academic, news and vector sources. Parallel agents share state and hand results to a consolidation step that checks for contradictions, with caching, cost tracking and multi-format exports around the core loop.

  • Python
  • LangGraph
  • LangChain
  • ChromaDB
  • Streamlit
Acquired

Enterprise risk platform

CERICO

Acquired by Dow Jones in March 2018.

A cloud platform for auditable third-party risk assessment. I led the functional architecture across questionnaires, scoring, workflow states, approvals, escalations, controls and reporting.

  • Functional architecture
  • Risk models
  • Audit evidence
In production

Conversational AI and operations

Real-estate intelligence platform

Built and deployed for two Dubai brokerages.

Connects conversational property discovery to leads, CRM, real-time reporting and financial workflows in one platform.

Architecture and scope

Architecture. A modular live-session client with audio streaming, a tool registry, Maps and Places grounding, shared state, persistence and analytics.

Operations. Chat-to-lead registration, CRM lifecycle, AI-assisted receipt scanning, milestone-based commission invoicing and tax invoices designed for UAE VAT.

Confidentiality. Proprietary client work; source code is not public.

  • Live audio
  • Tool calling
  • Google Maps
  • CRM workflows
In production

Applied machine learning

Trading and analytics systems

Prediction error below 1% after a correction algorithm I designed.

Market-data pipelines and ML services for prediction, correction, forecasting and trade analytics.

Role and contribution

Role. Head of Engineering Unit, supervising three quantitative analysts and two software engineers from requirements through validation.

Contribution. A new combination algorithm inside the correction engine, a layer that runs after stock-price prediction and before corrected prices feed order and signal calculations.

  • Python
  • TensorFlow
  • PostgreSQL
  • Flask
Published

Model safety research

Refusal control in Llama 3 8B

Accepted at NeurIPS 2025 and COLM 2026 interpretability workshops.

Mechanistic-interpretability research into category-specific refusal directions and steering them at inference time, with the Algoverse AI Research team.

  • LLM safety
  • Activation steering
  • Evaluation

// Research

AI safety and controllability

What refusal behaviour looks like inside a model, and how to steer it.

Accepted COLM 2026, Actionable Interpretability Workshop

From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions

Rishab Alagharu, Ishneet Sukhvinder Singh, Shaibi Shamsudeen, Zhen Wu, Ashwinee Panda

Poster NeurIPS 2025, Mechanistic Interpretability Workshop

What Do Refusal Tokens Learn? Fine-Grained Representations and Evidence for Downstream Steering

Presented as a workshop poster with the Algoverse AI Research team

// In development

Platform-level AI safety for children

A regulatory assurance concept that turns age-assurance and child-safety obligations into executable platform controls and verifiable evidence.

My working hypothesis: platforms need more than an age signal. They need to prove which rule applied, which control ran, whether it stayed effective, and when evidence must be refreshed or challenged.

Current stage
Competitive-gap analysis and customer validation

A concept in research and validation, not a finished or commercially available product.

  1. 1
    Age evidenceA provider or platform signal
  2. 2
    Assurance stateProvenance, level and freshness
  3. 3
    Policy ruleA decision the customer configures
  4. 4
    Executed controlThe platform action and its evidence
  5. 5
    Continuing assuranceExpiry, revocation, conflict and re-checks

// Career

From enterprise controls to AI systems

The thread through all of it is functional architecture: turning policy, risk, data and operating needs into systems people can use and verify.

Download the full resumePDF, 60 KB. Complete role history, dates and tooling.
2024
MSc Data Science, Collège de Paris
2013
MSc International Management, University of Liverpool
Certs
CCEP-I, ISO 37001 and ISO 9001 auditor training
  1. Jan 2024 to now

    Head of Engineering Unit, Proceedit Trading

    Lead systems architect for AI-driven trading engines and analytics services, leading three quantitative analysts and two software engineers. Designed the correction-engine algorithm that brought prediction error below 1%.

  2. May to Dec 2023

    Data Analyst, Avenir Arabia IT DMCC

    Cleaned and consolidated multi-source data, built Python, Power BI and Tableau dashboards, and defined KPIs to turn client process challenges into actionable insights.

  3. Jan 2021 to Apr 2023

    Head of Product Development and Compliance Services, Tadashie FZCO

    Led functional and product design for GRC applications across due diligence, conflicts of interest, anti-bribery controls, approvals and audit evidence. Improved process efficiency by 40%.

  4. Oct 2013 to Apr 2020

    Senior Compliance Advisor and Functional Systems Lead, Petrofac

    Led the functional architecture and product requirements for CERICO, defining risk models, questionnaires, workflow logic and approval controls. Built global compliance workflows with ERP integration.

  5. Jan 2008 to Oct 2013

    Quality Assurance Coordinator, Petrofac

    Designed an internal audit planning and findings-analysis tool and automated inspection coordination, reducing inspection costs by 30%, across major FEED and EPC projects.

Technology experience

Loop engineering
Agentic control loops, reason-act-observe cycles, reflection and self-correction, evaluator-optimizer loops, termination and fallback logic, human-in-the-loop approvals
Prompt engineering
System and role prompt design, context engineering, structured outputs, function and tool calling, prompt versioning and regression testing, injection defence
AI systems
Multi-agent orchestration, RAG, semantic search, vector databases, model routing, caching and token budgeting
Reliability and evaluation
Eval harnesses, guardrails, backtesting, observability and tracing, activation steering, cost and latency trade-offs
Engineering
Python, TypeScript, SQL, React, FastAPI, PostgreSQL, Supabase, TensorFlow, PyTorch
Delivery and tooling
AWS, Docker, CI/CD, Vercel, Power BI, Claude Code, Windsurf, Devin

The full list, with versions and dates, is in the resume.

// Let's talk

Hiring for AI systems architecture?

I'm looking for senior roles where engineering and control have to work together. Tell me a little about the role and I'll reply within a couple of days.

Opens your email app with this message filled in. Nothing is sent until you press send.