David RylandSenior AI Engineer • Full‑Stack Engineer
Summary

Senior FullStack Engineer and AI Engineer with senior‑level expertise delivering production‑grade, latency‑sensitive AI solutions. Led the design and launch of a real‑time AI voice agent platform on GCP, integrating Whisper, Gemini embeddings and distributed inference to achieve sub‑second response times, also built a multimodal retrieval engine serving millions of users with sub‑100/ms latency at Zillow. Proven track record of end‑to‑end architecture, from audio ingestion pipelines to LLM‑powered agents and vector search. Aiming to apply this experience to accelerate AI‑driven product innovation and operational excellence. 

Core Skills
AI & LLM Engineering: Real‑time AI voice systems (Whisper, Deepgram STT, ElevenLabs TTS, WebRTC, LiveKit) • Gemini 2.5/3 • Gemini Pro Vision • Vertex AI (Model Garden, Vector Search, Workbench, Pipelines) • LangGraph • LlamaIndex • Tool‑calling pipelines • Multimodal & multilingual embeddings • Cross‑lingual contrastive learning • RAG architectures • Retrieval grounding • Query rewriting • Attribute extraction • FAISS/pgvector • Redis vector serving • ONNX • Quantization • Triton • Drift detection • Hallucination monitoring • Auditability
Backend & Distributed Systems: Python (FastAPI, Django) • Node.js • Go • Nest.js • Java • Spring Boot • .NET Core • High‑throughput REST & streaming APIs • Async processing • WebSockets • FFmpeg audio ingestion • Microservices • Event‑driven systems • Autoscaling • Redis caching • PostgreSQL • Batch pipelines
Frontend Engineering: React • Next.js • Angular • TypeScript • Redux • High‑performance dashboards • Clinician‑facing UIs • Component systems • Accessibility standards
Cloud, DevOps & Infrastructure: AWS (EC2, S3, RDS, Lambda) • GCP (Vertex AI, BigQuery, Cloud Run, Cloud Functions, GKE, Pub/Sub, Cloud Storage, Cloud SQL, AlloyDB, IAM, VPC) • Azure • Distributed GPU training pipelines • Dataset generation workflows • Docker • GitHub Actions • CI/CD • Terraform • IaC • Cost‑efficient architecture design
Data, ML & Retrieval Systems: Databricks • Delta Lake • BigQuery analytics • ETL pipelines • Synthetic data generation • Hybrid search • Semantic ranking • Contextual grounding • Large‑scale batch inference
Media & Real Time Systems: FFmpeg • OpenCV • WebRTC • LiveKit • Twilio • Low‑latency audio/video processing • Real‑time streaming inference • Bidirectional audio APIs
Professional Experience
Striveworks LLC, Senior FullStack Engineer / AI Engineer
01/2025 – 06/2026 | Seattle WA
  • Architected and shipped a real‑time AI voice agent platform using Python, FastAPI, Django, LiveKit, WebRTC, and Twilio, deployed on GCP (Cloud Run, Cloud Functions, Pub/Sub, Cloud Storage) for scalable low‑latency inference.
  • Integrated Gemini 2.5 Pro and Vertex AI Embeddings into multimodal RAG pipelines combining audio, text, and metadata retrieval, improving grounding accuracy and reducing hallucinations.
  • Built streaming inference pipelines integrating Whisper/Deepgram STT, ElevenLabs TTS, and multimodal embeddings, delivering sub‑second transcription and response generation for interactive applications.
  • Designed interrupt‑driven agent logic with turn‑taking and barge‑in detection using an event‑driven state machine in FastAPI, implementing context‑aware state management for natural conversational flow across LLM‑powered agents.
  • Developed end‑to‑end audio ingestion and preprocessing pipelines using FFmpeg, segmentation, metadata extraction, and validation layers, ensuring reliable downstream inference for both real‑time and batch workloads.
  • Implemented multimodal AI pipelines combining audio, text, and metadata retrieval using RAG, vector search (FAISS/pgvector), and LLM tool‑calling, improving grounding accuracy and reducing hallucinations.
  • Built real‑time WebSocket APIs for bidirectional audio streaming, model inference, and agent orchestration, supporting concurrent sessions under strict latency budgets.
  • Implemented multi‑agent orchestration using LangGraph and LlamaIndex, enabling autonomous workflows, multi‑step reasoning, and tool‑calling pipelines for complex conversational and retrieval tasks.
  • Built batch inference pipelines for large‑scale evaluation, dataset generation, and offline scoring, enabling continuous improvement of speech and retrieval models.
  • Delivered reproducible deployments using Docker, GitHub Actions, AWS S3, and internal CI/CD pipelines, reducing deployment time and ensuring consistent releases.
  • Collaborated with speech scientists, clinicians, and product teams to validate outputs, refine system behavior, and improve prediction accuracy for clinician‑facing UIs through structured evaluation and user testing.
  • Leveraged AI coding tools (Claude Code, GitHub Copilot) to accelerate development, improve code quality, and streamline debugging across AI and full‑stack systems.
  • Mentored junior engineers through architecture reviews, pair programming, and AI workflow training, improving team velocity and raising engineering standards.
  • Led cross‑functional initiatives coordinating engineering, product, and research teams to deliver new AI features such as multimodal embedding services within tight deadlines.
  • Zillow, Senior FullStack Engineer / AI Engineer
    11/2022 – 12/2024 | Seattle, WA
  • Delivered end‑to‑end AI and full‑stack applications using React, Next.js, TypeScript, Python, FastAPI, Django, Node.js, Redis, PostgreSQL, and AWS, accelerating feature rollout and improving developer productivity across internal AI platforms.
  • Built and productionized multimodal, multilingual embedding‑based retrieval models serving 20M+ users with 10–100ms P99 latency and 10k+ QPS, deployed on GCP (GKE, Cloud Run, BigQuery, Vertex AI).
  • Designed and deployed cross‑lingual contrastive learning pipelines, synthetic data generators, and dataset scaling workflows on distributed AWS GPU clusters, increasing model coverage and training throughput.
  • Architected conversational search systems using query rewriting, attribute extraction, and RAG grounding across Next.js UIs and FastAPI/Django microservices, improving search relevance and user engagement.
  • Implemented multi‑agent orchestration with LangGraph and LlamaIndex, reducing manual query planning and cutting end‑to‑end response time through autonomous tool‑calling and multi‑step reasoning.
  • Built AI observability and governance frameworks using LangSmith, Langfuse, PostgreSQL, and custom telemetry dashboards, enabling drift detection, hallucination monitoring, prompt evaluation, and compliance‑aligned audit trails.
  • Integrated Gemini Pro Vision with Whisper and Deepgram speech‑to‑text, text‑to‑speech, and embedding models to create multimodal retrieval prototypes that improved query relevance and enabled new media‑based search
  • Led FAISS ANN index optimization, improving recall and reducing latency, and integrated updated vectors into a Redis‑backed serving layer for high‑throughput, low‑latency retrieval.
  • Owned the ML lifecycle including ONNX deployment, quantization, GPU/CPU optimization, Triton inference server integration, threshold tuning, feature pipelines, latency analysis, and BigQuery analytics surfaced via Next.js dashboards.
  • Built batch inference pipelines for large‑scale evaluation, synthetic data generation, and offline model scoring, enabling continuous improvement of retrieval and ranking models.
  • Conducted applied research on MEVI‑based generative retrieval, deploying supporting services on AWS Lambda and EC2 with Python, enabling scalable query generation and faster experimentation.
  • Banner Health, FullStack Developer
    08/2019 – 10/2022 | Phoenix, AZ
  • Built a clinical decision support platform enabling natural‑language search across treatment guidance and patient insights, improving clinician search efficiency.
  • Built a 1M+ document retrieval system using hybrid search, semantic ranking, vector retrieval, and contextual grounding, achieving a 22% accuracy improvement.
  • Developed workflow automation services for clinical summarization and compliance‑driven processing using Docker, GitHub Actions, and CI/CD pipelines, which accelerated report generation and ensured consistent compliance.
  • Designed multimodal pipelines for document intelligence, image understanding, and structured content extraction using FFmpeg and custom validators, increasing extraction accuracy and reducing manual review effort
  • Built scalable backend services (Python, FastAPI, Go, async processing, streaming APIs) supporting 5,000+ concurrent users with 35% latency reduction.
  • Developed clinician‑facing React, Angular, and TypeScript applications that met accessibility standards and were delivered through Docker‑based CI/CD pipelines, resulting in faster deployments and higher user satisfaction.
  • GoDaddy, Software Developer
    07/2016 – 07/2019 | Tempe, AZ
  • Built and optimized platform‑level services on AWS with Node.js, Nest.js, and Terraform IaC, improving deployment reliability and reducing downtime.
  • Maintained backend services using Java (Spring Boot) and .NET Core, improving API reliability and reducing latency.
  • Developed responsive React, Angular, and Next.js applications for SMB tools with TypeScript and Redux, enabling faster feature delivery and improving user satisfaction.
  • Created reusable UI component libraries with TypeScript and Storybook, streamlining front‑end development and reducing code duplication across projects.
  • Integrated frontend applications with backend services via REST, GraphQL, and Apollo, ensuring consistent data flow and decreasing integration errors.
  • Designed media-processing pipelines with Delta Lake, AWS Lambda, and FFmpeg, enabling automated video transcoding and reducing manual processing time.
  • Tuned Node.js services and AWS compute resources, reducing processing costs.
  • Education
    Bachelor of Science, Computer Science, Arizona State University
    08/2012 – 05/2016 | Tempe, AZ