AI platform engineer with 7+ years operating production systems on GCP and AWS. I ship LLM infrastructure that an engineering team depends on daily — two Model Context Protocol servers used in production by 10+ engineers, and a 50-user LangGraph assistant hardened with fail-closed data isolation, approval-gated writes, and quota-aware concurrency control. Underneath sits consumer-scale platform work: Kubernetes services sustaining 3,000+ RPS for 50M+ users, and a 35% quarter-over-quarter cut in critical incidents.
MuslimPro — 170M+ downloads, 50M active users · Singapore
Shipped 2 internal Model Context Protocol (MCP) servers now used in production by 10+ engineers, exposing 15+ read-only cloud-observability and BigQuery tools to AI assistants for incident triage, resource inventory, cost analysis, and plain-language data questions, with defence-in-depth read-only guarantees and per-query cost caps
Built and operate Habibi, an internal multi-user AI assistant serving 50+ users on GCP (Terraform VM, FastAPI + LangGraph + MongoDB + Qdrant), engineered for reliability: bounded-concurrency gate with exponential backoff against Vertex AI quota, fail-closed per-user data isolation, self-healing OAuth re-auth, approval-gated external writes with audit logging, Fernet-encrypted tokens, and a SELECT-only byte-capped BigQuery proxy
Built and own the analytics data platform on BigQuery — Pub/Sub event ingestion, scheduled ETL and batch jobs, and operational-database to warehouse sync feeding ~5TB of datasets, plus modelled reporting tables and dashboards consumed by business teams
Operated high-throughput microservices on GKE sustaining 3,000+ RPS during peak seasonal traffic, with autoscaling, Pub/Sub event-driven fault isolation, real-time WebSocket delivery, and production monitoring and alerting
Led production incident investigations and root-cause analysis, introducing monitoring and alerting improvements that reduced critical incidents 35% quarter-over-quarter
Designed a hardened production GKE environment for a fleet of 10–20 microservices: private cluster with disabled public endpoints, Workload Identity for pod-level IAM, custom VPC with NAT gateway, bastion-only control-plane access, and Kubernetes Gateway API load balancing
Built auditable CI/CD with Terraform plan/apply via GitHub Actions and Cloud Build using keyless OIDC (Workload Identity Federation), GCS state backend, and manual production approval gates across 3 GCP projects
Cut deployment lead time 80% with an import-aware selective deployment pipeline using static dependency-graph analysis — Cloud Functions deploys dropped from 15 minutes to under 3 minutes per commit
Orchestrated GDPR user deletion across ~20 distributed services using Google Workflows with compensation handlers, idempotency guarantees, and full audit trails
Built the Qalbox OTT streaming pipeline end-to-end for a 1,000+ title catalogue serving 2M+ monthly active viewers, connecting a Next.js admin CMS to Tencent VOD transcoding via Cloud Run functions and Firestore, with Widevine/FairPlay DRM and in-app-subscription entitlement across mobile, web, and Smart TV
Reduced API latency 40% on high-concurrency endpoints through Redis caching, connection pooling, and query optimisation across MongoDB and PostgreSQL
Mentored 3–5 engineers through code and design reviews, and authored technical writing on CI/CD and microservice patterns
Senior Software Engineer ·Sysco Labs
Dec 2020 – Dec 2022
Freight rate management — logistics and supply chain · Sri Lanka
Owned 3 freight-rate-management platforms end-to-end — FRDT, CFM, and SID — on a centralised team reconciling freight rates across 185 operating companies, streaming 100k+ rate records per day through AWS Kinesis into Java (Spring Boot) services so distribution costs stayed consistent and auditable
Built 5+ distributed backend microservices in Java (Spring Boot) on AWS ECS at 99.9% uptime, with real-time pipelines on AWS Kinesis and Lambda
Established Terraform IaC and automated CI/CD that cut deployment time 70% while improving release stability, and improved critical API performance 60% via PostgreSQL indexing, query tuning, and read replicas
Implemented OAuth 2.0 across 2 identity providers (Amazon Cognito and Azure AD), and mentored engineers on microservice design and cloud-native principles
Software Engineer ·DirectFN
May 2019 – Dec 2020
Financial technology — real-time trading · Sri Lanka
Developed secure backend services for a real-time financial trading platform: encrypted file processing with token-based access control, audit logging, Docker containerisation, and production monitoring
Selected work
platform-mcp — Read-only MCP server turning a GCP project into an autonomous SRE surface — 15 observability tools across logs, metrics, errors, cost, recommenders, and inventory, with defence-in-depth read-only enforcement.
Limbiq — Neurotransmitter-inspired adaptive memory layer for LLMs — FAISS vector search, self-attention encoder, and knowledge-graph extraction. Published on PyPI.
marsClaw — Multi-channel LLM chat agent on the Claude Agent SDK and Gemini, with its own MCP server for Gmail, Calendar, and Drive tools, and a hardened security model.
FleetOS — Distributed humanoid-robot fleet platform — 17 Go services over gRPC and Kafka, Temporal workflows, Kubeflow GPU training, and a Spark/ClickHouse analytics path.
Kithly — Group video card product shipped end-to-end — React Native (Expo) app on iOS and Android, Go API, multi-stage FFmpeg pipeline, HLS transcoding, and S3/CloudFront delivery.