I’m currently working as a AI Engineer at iOPTIME PVT LTD. I design, build and ship generative AI systems, from prototyping to production deployments. My work covers Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), agentic and multi-agent orchestration (LangGraph), and the MLOps practices required to make those systems reliable and maintainable in production.
I develop retrieval-augmented systems leveraging vector databases, integrating them with large language models to enable context-aware, knowledge-grounded responses. I architect multi-step agent pipelines that decompose complex tasks into modular, deterministic subtasks for robust execution. In production, I implement FastAPI-based model serving and containerized deployments, ensuring scalable, reliable, and production-ready ML services. On the production side I build FastAPI-backed model services, containerize them with Docker, and wire them into CI/CD and monitoring pipelines so teams can safely push updates and operate at scale.
Beyond text, I have practical experience with real-time computer vision (SOTA segmentation models optimized for low-latency inference) and audio pipelines , including speech-to-text with OpenAI’s Whisper and speaker diarization for multi-speaker transcripts.
Contact: aashirali619@gmail.com
- Generative models & LLMs: prompt engineering, fine-tuning, instruction tuning, evaluation and calibration
- Retrieval & RAG: embeddings, vector databases (QDrant/ChromaDb), retrieval pipelines, and latency-aware design
- Agentic & multi-agent workflows: LangGraph orchestration, planning, and robust error handling between agents
- MLOps & Serving: FastAPI-based inference services, Docker containerization, CI/CD, observability and cost-aware deployment
- Computer Vision & Audio: segmentation model optimization for edge devices, Whisper-based Speech-To-Text, speaker diarization and transcript processing
- Designed and deployed multiple RAG systems integrating vector indexes with LLMs for knowledge-grounded generation and Q&A.
- Built multi-agent orchestration using LangGraph to decompose complex workflows into reliable agent sub-tasks, with clear state passing and retry policies.
- Packaged model services with FastAPI, containerized them using Docker, and integrated with CI/CD pipelines and monitoring for safe production operations.
- Implemented real-time segmentation models optimized for edge inference (pruning/quantization and runtime tuning).
- Built audio pipelines using Whisper for STT and pyannote-like approaches for speaker diarization, producing clean multi-speaker transcripts and metadata.





