
Production LLM agent runtime with multi-step tool use, rate limiting, idempotency, audit logging, token metering, and OpenTelemetry tracing.
Production RAG pipeline with BM25 + dense retrieval, cross-encoder reranking, RAGAS evaluation suite, and Prometheus metrics for faithfulness and latency monitoring.
Benchmarks comparing Full FT, LoRA, and QLoRA across GPU memory, training cost, inference latency, and task performance on instruction-following datasets.
Centralized LLM gateway, per-tenant policy enforcement, semantic caching, prompt injection detection, cost tracking per model and team, and unified observability across OpenAI and Anthropic.
End-to-end MLOps platform with experiment tracking, registry-based model promotion, canary serving with traffic splitting, statistical drift detection, and automated retraining triggers.
More projects available on GitHub