Profile picture
Manoj Kumar Nagabandi
Email
knownbymanoj@gmail.com
Location
Italy (Searching for relocation to Sweden)
GitHub
Demos of my work
Work Experience

Full Stack Data Scientist / AI Engineer

SIPA SpA (Zoppas Industries Group)
09/2023 – present | Italy

Sole AI engineer responsible for designing and productionizing the ECHO Platform, an enterprise AI assistant ecosystem embedded inside SIPA's industrial data platform serving 3,000+ internal staff and customers globally. The platform integrates ERP, CRM, PLM, ticketing, and IoT data streams, exposing them to end users through three specialized AI agents:

Echo Platform Assistant

  • Built an agentic real-time query system connecting to 25+ live REST APIs (machines, orders, offers, production data, alarms) using a hierarchical intent classification architecture (root → child intent → parameter extraction → parallel function calling → LLM-based answer synthesis).
  • Integrated directly into the Echo Platform back-end via WebSockets and JavaScript front-end, with full RBAC enforcement to ensure data access security at scale.
  • Implemented a real-time user feedback collection, Langfuse observability dashboards, and continuous prompt refinement via RAG evaluation loops.
  • Manuals Assistant (Technical Documentation RAG)

  • Designed an advanced hybrid RAG retrieval pipeline for industrial PDF manuals (maintenance, installation, operational), parsing hierarchical document structures via Table of Contents, storing chapter/sub-chapter hierarchies in PostgreSQL with dynamic linking to pgvector embeddings (OpenAI text-embedding-3-large, dim=256) stored in AWS S3.
  • Implemented hybrid search combining Full-Text Search (FTS) for keyword precision and vector semantic search, fused using Reciprocal Rank Fusion (RRF), followed by an LLM-based re-ranking/filtering stage and extractive summarization with highlighted citations.
  • Applied query rewriting using document ToC context and conversation history to handle multi-turn dialogue and ambiguous questions.
  • Help Desk Assistant (Ticket Intelligence)

  • Built a full data pipeline for historical service ticket processing: PII removal using a domain-specific technical glossary, structured metadata extraction (alarms, machine models, issue types), and embedding-based indexing for semantic retrieval.
  • Designed a Intelligent search with dual-mode retrieval system: deterministic query mode (structured ticket search by parameters) and RAG semantic mode (hybrid retrieval with deterministic filters), grounded with cited ticket references for traceability.
  • Cross-Platform LLMOps

  • Designed and implemented RAG evaluation frameworks (promptfoo) across all three assistants, benchmarking retrieval precision, task completion, tool selection, answer faithfulness, and contextual relevance using multiple metrics via Langfuse.
  • Currently architecting inter-agent handoff flows using PydanticAI to unify all three assistants into a single orchestrated multi-agent system with delegation and context propagation.
  • Skills
    Programming Languages

    Python (advanced) · SQL (proficient) · R · C

    AI / ML Frameworks

    PyTorch · TensorFlow · Keras · Scikit-learn · XGBoost · MLflow · PySpark

    Generative AI & LLMs

    OpenAI API · (Open source) HuggingFace Transformers · LangChain · LangGraph · LlamaIndex · PydanticAI Agents · MCP servers · LlamaIndex

    RAG & Search

    Hybrid retrieval (FTS + vector search) · Reciprocal Rank Fusion (RRF) · pgvector · Hierarchical PDF parsing · Query rewriting · Metadata filtering · Semantic chunking · Different Vector DB benchmarking

    Agentic & LLMOps
    • PydanticAI multi-agent orchestration · Intent classification · Inter-agent handoffs · Langfuse (observability, evaluation, prompt refinement) · RAG benchmarking · A/B evaluation · MCP Servers · Promptfoo evaluation
    Cloud Platforms

    AWS (S3, EC2, IoT, SageMaker) · GCP · Microsoft Azure

    Data Engineering & MLOps

    ETL pipeline design · REST API development (Django) · WebSockets · Docker · CI/CD · MLflow · Feature engineering · RBAC integration

    Databases & Storage

    PostgreSQL · pgvector · MongoDB · MySQL · AWS S3

    Other Tools
    • Jupyter Notebook · PowerBI · Git · PyCharm · Chainlit · Pdfminer
    Education

    Master's in Data Science

    University of Padova
    09/2021 – 09/2023 | Padova, Italy

    Specialized in Machine Learning, Deep Learning, NLP, Statistical Learning, and Big Data systems. Hands-on laboratory program covering model evaluation, fairness, and responsible AI. Familiar with legal, ethical, and business dimensions of data-driven products.

    Relevant Coursework: Machine Learning, Deep Learning, NLP, Statistics, Computer Vision, Big Data (Spark/MapReduce), Information Systems, Business Economics & Financial Data

    Bachelor's in Engineering Sciences

    Univeristy of Rome Tor Vergata
    09/2017 – 04/2021 | Rome, Italy

    Engineering foundations: network architectures, cryptography, algorithms, distributed systems, and systems design.

    Data Science Intern

    SIPA SpA⁠
    03/2023 – 08/2023 | Vittorio Veneto, Italia
  • Developed a Multiple Time Series Forecasting solution for 4,200 industrial machines simultaneously, applying XGBoost and Transformer-based architectures (LSTM, ARIMA) to predict machine alarm events and production cycle timing enabling proactive maintenance scheduling.
  • Applied unsupervised ML (clustering, deduplication algorithms) to clean and normalize large-scale industrial datasets, improving downstream model accuracy and data integrity.
  • Validated a newly built AWS IoT cloud architecture, supporting the company's infrastructure migration with data pipeline testing and integration checks.
  • Machine Learning Intern

    iNeuron⁠
    04/2021 – 09/2021 | Bengaluru, India
  • Built and deployed an end-to-end Forest Cover Type Classification pipeline on AWS (S3, EC2, SageMaker), achieving an AUC of 0.975 on test data and reducing operational field survey costs by ~90%.
  • Designed reusable ML pipelines covering data ingestion, feature engineering, model training, validation, and production deployment following CI/CD best practices.
  • COURSE WORK
  • MACHINE LEARNING
  • DEEP LEARNING
  • STATISTICS
  • FUNDAMENTALS OF INFORMATION SYSTEMS
  • Interests
    Gym
    Learning new languages, and cultures
    Learning new technologies to work more efficient and be better at my job
    Personal Projects

    Agentic FleetOps Copilot⁠

    Conversational Fleet Intelligence
    05/2026 – 05/2026

    Built a conversational fleet intelligence system that predicts truck component failures, explains risk with SHAP, and supports maintenance decisions through an agentic chat interface backed by FastAPI services. The project combines PydanticAI, XGBoost, Langfuse, and a Scania-inspired fleet simulator, using real predictive-maintenance patterns from the Component X dataset.

    Applied ARIMA, Generalized Additive Regression, CART, and Gradient Boosting to forecast BTCUSD price as part of a Business Economics & Financial Data course. Demonstrates financial data modeling and time series.

    Binary classification model to predict customer churn using feature selection, class balancing, and ensemble methods (Random Forest, Gradient Boosting). This project is directly applicable to customer analytics and lifecycle modeling.

    Benchmarked model agnostic feature selection techniques across SVM, Perceptron, Decision Trees, and ensemble models on varied datasets, demonstrating rigorous statistical experimentation and ML methodology.

    As part of Big Data Course used Map Reduce implementation using RDD (Resilient Distributed Dataset) for analyzing the huge scale datasets

    Using R Language I have done analysis and removed outliers and used statistical models like Poison GLM and forward and backward selection techniques to find optimal features by AIC and BIC.

    Languages
    English , Hindi, Telugu

    Full working Proficiency

    Italian

    Working Proficiency