David RylandSenior AI Engineer • Full‑Stack Engineer
Summary
Senior FullStack Engineer and AI Engineer with senior‑level expertise delivering production‑grade, latency‑sensitive AI solutions. Led the design and launch of a real‑time AI voice agent platform on GCP, integrating Whisper, Gemini embeddings and distributed inference to achieve sub‑second response times, also built a multimodal retrieval engine serving millions of users with sub‑100/ms latency at Zillow. Proven track record of end‑to‑end architecture, from audio ingestion pipelines to LLM‑powered agents and vector search. Aiming to apply this experience to accelerate AI‑driven product innovation and operational excellence.
Core Skills
AI & LLM Engineering: Real‑time AI voice systems (Whisper, Deepgram STT, ElevenLabs TTS, WebRTC, LiveKit) | Gemini 2.5/3 | Gemini Pro Vision | Vertex AI (Model Garden, Vector Search, Workbench, Pipelines) | LangGraph | LlamaIndex | Tool‑calling pipelines | Multimodal & multilingual embeddings | Cross‑lingual contrastive learning | RAG architectures | Retrieval grounding | Query rewriting | Attribute extraction | FAISS/pgvector | Redis vector serving | ONNX | Quantization | Triton | Drift detection | Hallucination monitoring | Auditability
Backend & Distributed Systems: Python (FastAPI, Django) | Node.js | Go | Nest.js | Java | Spring Boot | .NET Core | High‑throughput REST & streaming APIs | Async processing | WebSockets | FFmpeg audio ingestion | Microservices | Event‑driven systems | Autoscaling | Redis caching | PostgreSQL | Batch pipelines
Frontend Engineering: React | Next.js | Angular | TypeScript | Redux | High‑performance dashboards | Clinician‑facing UIs | Component systems | Accessibility standards
Cloud, DevOps & Infrastructure: AWS (EC2, S3, RDS, Lambda) | GCP (Vertex AI, BigQuery, Cloud Run, Cloud Functions, GKE, Pub/Sub, Cloud Storage, Cloud SQL, AlloyDB, IAM, VPC) | Azure | Distributed GPU training pipelines | Dataset generation workflows | Docker | GitHub Actions | CI/CD | Terraform | IaC | Cost‑efficient architecture design
Data, ML & Retrieval Systems: Databricks | Delta Lake | BigQuery analytics | ETL pipelines | Synthetic data generation | Hybrid search | Semantic ranking | Contextual grounding | Large‑scale batch inference
Media & Real Time Systems: FFmpeg | OpenCV | WebRTC | LiveKit | Twilio | Low‑latency audio/video processing | Real‑time streaming inference | Bidirectional audio APIs
Professional Experience
Striveworks LLC, Senior FullStack Engineer / AI Engineer
01/2025 – 06/2026 | Remote
  • •Architected and shipped a real-time AI voice agent platform using Python, FastAPI, Django, LiveKit, WebRTC, and Twilio, deployed on GCP (Cloud Run, Cloud Functions, Pub/Sub, Cloud Storage) for scalable low-latency inference; partnered with clinical and product stakeholders to translate functional requirements into production agentic solutions.
  • •Integrated Gemini 2.5 Pro and Vertex AI Embeddings into multimodal RAG pipelines combining audio, text, and metadata retrieval, improving grounding accuracy and reducing hallucinations through prompt/context engineering and retrieval grounding.
  • •Built streaming inference pipelines integrating Whisper/Deepgram STT, ElevenLabs TTS, and multimodal embeddings, delivering sub-second transcription and response generation for interactive applications.
  • •Designed interrupt-driven agent logic with turn-taking and barge-in detection using an event-driven state machine in FastAPI, implementing context-aware memory/state management and human-in-the-loop controls for natural conversational flow across LLM-powered agents.
  • •Developed end-to-end audio ingestion and preprocessing pipelines using FFmpeg, segmentation, metadata extraction, and validation layers, ensuring reliable downstream inference for both real-time and batch workloads across structured and unstructured data.
  • •Implemented multimodal AI pipelines combining audio, text, and metadata retrieval using RAG, vector search (FAISS/pgvector), and LLM tool/function calling, improving grounding accuracy and reducing hallucinations.
  • •Built real-time WebSocket APIs and API clients for bidirectional audio streaming, model inference, and agent orchestration, supporting concurrent sessions under strict latency budgets and integrating with enterprise systems and third-party services.
  • •Implemented multi-agent orchestration using LangGraph and LlamaIndex, enabling autonomous workflows, multi-step reasoning, tool-calling pipelines, and failure-aware execution for complex conversational and retrieval tasks.
  • •Built batch inference pipelines for large-scale evaluation, dataset generation, and offline scoring, enabling continuous improvement of speech and retrieval models with clear quality measurement and observability.
  • •Delivered reproducible deployments using Docker, GitHub Actions, AWS S3, and internal CI/CD pipelines (including automated tests), reducing deployment time and ensuring consistent releases.
  • •Collaborated with speech scientists, clinicians, and product teams to validate outputs, refine system behavior, and improve prediction accuracy for clinician-facing UIs through structured evaluation, user testing, and iterative troubleshooting.
  • •Leveraged AI coding tools (Claude Code, GitHub Copilot) daily across the software development lifecycle to accelerate development, improve code quality, and streamline debugging of AI and full-stack systems.
  • •Mentored junior engineers through architecture reviews, pair programming, and AI workflow training, improving team velocity while operating in ambiguous, fast-paced environments.
  • 1 / 2
  • •Led cross-functional initiatives coordinating engineering, product, and research teams to deliver new AI features such as multimodal embedding services within tight deadlines, converting requirements into production value.
  • Zillow, Senior FullStack Engineer / AI Engineer
    11/2022 – 12/2024 | Seattle WA
  • •Delivered end-to-end AI and full-stack applications using React, Next.js, TypeScript, Python, FastAPI, Django, Node.js, Redis, PostgreSQL, and AWS, accelerating feature rollout across internal AI platforms while translating business needs into technical designs.
  • •Built and productionized multimodal, multilingual embedding-based retrieval models serving 20M+ users with 10–100 ms P99 latency and 10k+ QPS, deployed on GCP (GKE, Cloud Run, BigQuery, Vertex AI); owned discovery through production integration and adoption.
  • •Designed and deployed cross-lingual contrastive learning pipelines, synthetic data generators, and dataset scaling workflows on distributed AWS GPU clusters, increasing model coverage and training throughput for large-scale structured and unstructured data.
  • •Architected conversational search systems using query rewriting, attribute extraction, and RAG grounding across Next.js UIs and FastAPI/Django microservices, improving search relevance and user engagement through agentic workflows.
  • •Implemented multi-agent orchestration with LangGraph and LlamaIndex, reducing manual query planning and cutting end-to-end response time through autonomous tool-calling, multi-step reasoning, and state management.
  • •Built AI observability and governance frameworks using LangSmith, Langfuse, PostgreSQL, and custom telemetry dashboards, enabling drift detection, hallucination monitoring, prompt evaluation, execution tracing, safety controls, and compliance-aligned audit trails.
  • •Integrated Gemini Pro Vision with Whisper and Deepgram speech-to-text, text-to-speech, and embedding models to create multimodal retrieval prototypes that improved query relevance and enabled new media-based search capabilities.
  • •Led FAISS ANN index optimization, improving recall and reducing latency, and integrated updated vectors into a Redis-backed serving layer for high-throughput, low-latency retrieval; applied strong API, database, and access-control patterns.
  • •Owned the ML lifecycle including ONNX deployment, quantization, GPU/CPU optimization, Triton inference server integration, threshold tuning, feature pipelines, latency analysis, and BigQuery analytics surfaced via Next.js dashboards.
  • •Built batch inference pipelines for large-scale evaluation, synthetic data generation, and offline model scoring, enabling continuous improvement of retrieval and ranking models with measurable quality metrics.
  • •Conducted applied research on MEVI-based generative retrieval, deploying supporting services on AWS Lambda and EC2 with Python, enabling scalable query generation and faster experimentation.
  • Banner Health, FullStack Developer
    08/2019 – 10/2022 | Phoenix, AZ
  • •Built a clinical decision support platform enabling natural-language search across treatment guidance and patient insights, improving clinician search efficiency; embedded with clinical and operations teams to discover high-value problems and deliver AI-led automation with measurable adoption impact (healthcare domain experience).
  • •Built a 1M+ document retrieval system using hybrid search, semantic ranking, vector retrieval, and contextual grounding, achieving a 22% accuracy improvement; took solutions from discovery and MVP through production integration and ongoing use.
  • •Developed workflow automation services for clinical summarization and compliance-driven processing using Docker, GitHub Actions, and CI/CD pipelines, which accelerated report generation and ensured consistent compliance; demonstrated end-to-end ownership of AI workflow automation.
  • •Designed multimodal pipelines for document intelligence, image understanding, and structured content extraction using FFmpeg and custom validators, increasing extraction accuracy and reducing manual review effort; handled structured and unstructured clinical data at scale.
  • •Built scalable backend services (Python, FastAPI, Go, async processing, streaming APIs) supporting 5,000+ concurrent users with 35% latency reduction; applied strong API, database, authentication, and deployment fundamentals.
  • •Developed clinician-facing React, Angular, and TypeScript applications that met accessibility standards and were delivered through Docker-based CI/CD pipelines, resulting in faster deployments and higher user satisfaction; collaborated closely with business stakeholders on requirements, validation, and adoption.
  • GoDaddy, Software Developer
    07/2016 – 07/2019 | Tempe, AZ
  • •Built and optimized platform-level services on AWS with Node.js, Nest.js, and Terraform IaC, improving deployment reliability and reducing downtime; established early patterns for cloud-native, production-grade systems.
  • •Maintained backend services using Java (Spring Boot) and .NET Core, improving API reliability and reducing latency; gained hands-on experience with enterprise-style service integration and authentication patterns.
  • •Developed responsive React, Angular, and Next.js applications for SMB tools with TypeScript and Redux, enabling faster feature delivery and improving user satisfaction.
  • •Created reusable UI component libraries with TypeScript and Storybook, streamlining front-end development and reducing code duplication across projects.
  • •Integrated frontend applications with backend services via REST, GraphQL, and Apollo, ensuring consistent data flow and decreasing integration errors—foundational experience with APIs, third-party services, and data platforms.
  • •Designed media-processing pipelines with Delta Lake, AWS Lambda, and FFmpeg, enabling automated video transcoding and reducing manual processing time; early exposure to ETL-style pipelines and unstructured data handling.
  • •Tuned Node.js services and AWS compute resources, reducing processing costs while maintaining production reliability and observability.
  • Education
    Bachelor of Science, Computer Science, Arizona State University
    08/2012 – 05/2016 | Tempe, AZ
    2 / 2