Joseph CasilangSenior AI Software Engineer
Profile
Senior AI Software Engineer specializing in high performance, HIPAA compliant AI systems, multi agent orchestration, and enterprise scale multimodal pipelines. Architected production LLM platforms with sub 120ms latency, 99.995% uptime, and 40% cost reduction across healthcare and retail environments. Designed hybrid semantic search engines, dynamic model routing frameworks, and GPU autoscaling clusters serving 50,000+ users and processing 2M+ daily clinical and retail events. Led Zero Trust AI security initiatives, SOC2/NIST compliance, and multi cloud deployments across AWS, Azure, and GCP. Known for delivering measurable outcomes including 35% accuracy improvement, 70% manual workload reduction, and 25% faster feature delivery.
Skills
AI & Machine Learning: LLMs (GPT 4o, Claude 3.5, Gemini), multi agent systems (MCP, LangGraph), RAG pipelines, hybrid semantic search (FAISS, pgvector, Pinecone), dynamic model routing, multimodal embeddings, GPU autoscaling, ONNX optimization, quantization, hallucination reduction frameworks, evaluation pipelines, synthetic data generation
Backend & Platform Engineering: Python (FastAPI, Django), Node.js, Go, Java, microservices, event driven architecture (Kafka, EventBridge), high throughput REST/GraphQL APIs, WebSockets, distributed systems, Zero Trust backend design, RBAC, encryption, audit logging
Frontend: React, Next.js, JavaScript, TypeScript
Real Time & Voice Systems: Whisper, Deepgram, ElevenLabs, LiveKit, WebRTC, low latency streaming pipelines, bidirectional audio APIs, multimodal STT → LLM → TTS agents
Cloud & Infrastructure: AWS (EKS, ECS, Fargate, Lambda, Bedrock, SageMaker), Azure (AKS, Azure OpenAI, Document Intelligence), GCP (GKE, Cloud Run, Vertex AI), Kubernetes, Docker, Terraform, CI/CD (GitHub Actions), GPU autoscaling clusters, multi cloud orchestration, spot instance optimization
Data Systems: PostgreSQL, MongoDB, Redis, Supabase, Snowflake, DynamoDB, ETL/ELT pipelines, dbt, Dagster, PySpark, 1M+ daily event ingestion
Professional Experience
Roger Healthcare, Senior AI Software Engineer
05/2024 – 04/2026 | Fairfield, CA
  • Architected a multi agent clinical AI platform using LangGraph + MCP, enabling autonomous documentation, triage, and care plan generation with 70% reduction in clinician manual workload.
  • Built a hybrid semantic search engine combining Pinecone + pgvector + Azure Document Intelligence, delivering 35% accuracy improvement and sub 120ms retrieval latency.
  • Designed a multi cloud LLM routing layer across Azure OpenAI, AWS Bedrock, and GCP Vertex AI, reducing inference cost by 40% while maintaining 99.995% uptime.
  • Implemented GPU autoscaling clusters on EKS + GKE for real time transcription (Whisper/Deepgram) and multimodal inference, supporting 50,000+ users and 2M+ daily events.
  • Delivered Zero Trust AI security with OIDC federation, KMS encryption, SOC2/NIST controls, and PHI safe pipelines, passing all compliance audits with zero findings.
  • Built real time STT → LLM → TTS agents using LiveKit + WebRTC, reducing average response time by 45% and improving clinician satisfaction scores.
  • Led CI/CD modernization using GitHub Actions + Sigstore signing, achieving 100% deployment error elimination and 25% faster release cycles.
  • NextGen Healthcare, Senior Full Stack Engineer
    07/2021 – 04/2024 | Atlanta, GA
  • Built enterprise grade RAG assistants for clinical workflows using LangChain + hybrid search, improving documentation accuracy by 32% and reducing follow up edits by 60%.
  • Designed multimodal classification pipelines using LoRA fine tuned models, improving entity extraction accuracy by 28% across 1M+ clinical records.
  • Architected a distributed scheduling engine using Kafka + FastAPI + Redis, processing 1M+ daily scheduling events with 99.99% reliability.
  • Implemented multi cloud inference routing (AWS Bedrock + Azure OpenAI), reducing latency by 38% and cutting inference cost by 22%.
  • Built clinician facing React/Next.js dashboards with real time WebSocket updates, improving workflow visibility and reducing task turnaround time by 35%.
  • Migrated 500+ enterprise clients to new clinical data models with zero downtime, improving data consistency and reducing support tickets by 40%.
  • Carmatec, Full Stack Engineer
    06/2018 – 06/2021 | Los Angeles, CA
  • Built retail + healthcare AI automation systems serving 50,000+ monthly users, integrating RAG pipelines and multimodal search into customer facing dashboards.
  • Designed event driven retail inventory pipelines using DynamoDB + Kafka, achieving sub 200ms update latency across global storefronts.
  • Migrated legacy retail systems to AWS EKS + Docker, reducing infrastructure cost by 30% and improving deployment reliability.
  • Built internal analytics dashboards using React + Next.js + PostgreSQL, reducing reporting turnaround time by 50%.
  • Mentored junior engineers on AI workflow design, improving team velocity by 20%.
  • Frontend Developer
    05/2016 – 06/2018 | Los Angeles, CA
  • Built high performance React/TypeScript dashboards with real time WebSocket updates, reducing clinician decision latency by 25%.
  • Designed reusable component libraries that reduced UI development time by 40%.
  • Implemented performance optimizations (lazy loading, bundle splitting) that cut load times by 35% across healthcare and retail apps.
  • Education
    University of Santo Tomas, Bachelor's Degree
    08/2012 – 05/2016