Projects
Production systems, live demos, and open source.
Production AI systems
enterprise scaleSystems designed and operated at enterprise scale inside large engineering organizations. Described at a high level; source is proprietary.
LLM Evaluation Framework
Designed and operated the evaluation framework for generative AI products in production. Defined methodology across retrieval quality (Recall@k, Precision@k, MRR) and generation quality (groundedness, faithfulness), built on Databricks and Spark and adopted as standard practice across AI product teams.
RAG Pipeline Infrastructure
Architected production RAG pipelines connecting LLMs to enterprise knowledge sources. Designed embedding workflows, vector search infrastructure, and retrieval optimization for high-stakes enterprise workloads where hallucination risk is unacceptable.
Agentic Workflow Platform
Designed and deployed agentic AI systems in production: multi-turn orchestration, function calling, and tool-use patterns at enterprise scale across AWS and Azure. Built the infrastructure enabling LLMs to execute multi-step tasks autonomously against internal systems.
Internal Developer Platform
Led the engineering platform serving hundreds of developers across build, ship, and operate workflows. Designed secure-by-default CI/CD pipelines, containerization strategies, and IaC tooling that drove the shift from monolithic deployments to microservices.
Platform reliability & operations
production disciplineSystems and practices ensuring platform-critical infrastructure stays reliable and incident response closes the loop.
Staged Rollout Framework (MAP)
Architected and operated a three-stage onboarding framework (NPE → Canary → Production) governing how Business Units adopt the Moody's Analytics Platform. Clear readiness gates and evaluation criteria ensure platform changes go through structured assessment before affecting all downstream applications, preventing cascade failures.
Observability & Incident Management
Guided enterprise engineering and SRE teams onto Datadog observability stack with PagerDuty integrations. Coordinated incident response and post-mortem ownership with platform core and SRE teams, ensuring a platform-level outage doesn't cascade to every dependent application.
Live demos
running nowRunning systems with public source. You can interact with them or read the code.
Latino Canon 294 titles live rag llm-eval mcp workers-ai
Search for Latino-led and Latino-focused films and series, in English or Spanish. Hybrid retrieval (BM25 over D1 FTS5 plus bge-m3 in Vectorize, fused with reciprocal rank fusion) across four Cloudflare Workers. A 97-query golden set runs after every deploy (recall@5 0.818), a versioned 70B groundedness judge gates every generated note (auto-approval 33% → 87% after error analysis), and a remote MCP server lets Claude search the catalog. Connect it in claude.ai with latino-canon-mcp.ai-builders-studio-latinx.workers.dev/mcp. $0/month on the free tier, with a runbook built from real production incidents.
edge-ai-agent-lab live mcp workers-ai
Live Cloudflare Worker implementing the Model Context Protocol (MCP). Exposes three tools (time_now, worker_info, and echo) demonstrating MCP server patterns, Workers AI binding, and edge AI deployment. A reference implementation for agentic tool-use architecture.
TrustClaw 28 deployments live claude jfrog vercel
Autonomous email-summarization agent using Claude Haiku with function calling for tool-based auto-reply. Demonstrates Claude tool-use patterns, serverless deployment to Vercel, and supply-chain auditability through JFrog Artifactory proxying 900+ packages. Includes 5-point DX friction report with actionable fixes.
AI-Vic: conversational AI on the edge LLM-judged live rag agents llm-eval workers-ai
The portfolio assistant on this site. RAG over a Cloudflare Vectorize corpus plus 8 agentic tools, wrapped in a production LLM-as-judge evaluation layer: a small model scores every reply for relevance and groundedness — nightly against a fixed regression suite and on every visitor 👍/👎 — with results cached by (query, reply) so continuous evaluation adds near-zero cost. The judge scored everything 5/5 for two weeks until per-question hard-case rubrics gave it something to fail on; it now carries matched guards against both fabrication and over-cautious deflection. Human ratings sit beside the machine verdict as a labelled human-vs-machine dataset. Free-tier Cloudflare Worker on Llama 3.3 70B; hardened through a live neuron-quota incident that was quantitatively root-caused and fixed the same night. 280 tests, red-teamed with zero findings. The methodology is public in the ai-vic-eval repo below. Open the chat widget to try it.
AI-Vic Fine-Tune: the LoRA Adapter I Didn't Ship negative result lora fine-tuning huggingface
Trained a LoRA adapter (rank 8, 16.8 MB) on AI-Vic's own LLM-judged high-scoring replies, across two base models, and A/B tested every version against the un-adapted base with the same judge that scores production traffic. Every adapter regressed: the judge-filtered training data skewed toward confident assertion, so the model learned to always assert and fabricated on questions the corpus doesn't cover. Shipped the base model plus RAG instead. The adapter and full model card stay on Hugging Face as the documented artifact; the write-up on Medium is the point. Reliability by default, impressiveness by choice.
Open source
github.com/ramirez-ai-labsPublic repos covering RAG evaluation, LLM benchmarking, agentic systems, and AI foundations.
latino-canon v1.6.0 rag eval cloudflare
Source for the Latino Canon live demo above: a pnpm/Turborepo monorepo with four Workers, a shared eval package (recall@k, MRR, nDCG@10, groundedness judge, MCP tool-selection eval), eval-checked deploys, and about 330 tests. The README maps every skill to the code and PR that shows it.
ai-vic-eval new llm-as-judge rag methodology
The LLM-as-judge evaluation methodology behind AI-Vic: the judge system prompt with the rationale for every clause, the hard-case rubrics, the training-export filters, a sample nightly scorecard, and the full fine-tuning negative-result write-up. Published so the eval work is verifiable to a hiring manager without exposing the private chatbot repo.
RAG Evaluation Lab rag eval python
Part 1 of a 3-part learning path that swaps the evaluation stack while keeping the same Recall@k / Precision@k / MRR / groundedness core. Fully offline and beginner-friendly: synthetic datasets, embeddings, vector search, keyword vs TF-IDF retrieval, no API dependencies. The most-starred repo in the portfolio, used directly in Techqueria workshops.
GCP RAG Evaluation Lab rag eval vertex-ai
Part 2 of the series: production-grade RAG evaluation with Vertex AI embeddings, BigQuery persistence, and Gemini Vision multimodal extraction. Same metrics core, cloud-native stack.
Snowflake RAG Evaluation Lab rag eval snowflake
Part 3: a Forward Deployed Engineer's RAG evaluation playbook on Snowflake Cortex embeddings and native compute, no vendor lock-in. Decision frameworks for converting customer constraints (timeline, budget, compliance) into a retriever recommendation, plus regression tracking.
AI Operating System (AI-OS) agents mcp langgraph
Enterprise AI operating system orchestrating multi-agent workflows with Claude as the default LLM provider. Features a CI-gated evaluation harness running on every commit, a tool-loop safety circuit breaker preventing runaway agent behavior, and strict output validation requiring cited sources or failing at parse time. Demonstrates production patterns for Claude tool-use and agentic safety.
Chatbot Evaluation System llm-as-judge ci
Black-box pairwise evaluation baseline for comparing chatbot versions using LLM-as-judge. Supports fake and edge judge paths, audit metadata, and CI integration, with a roadmap toward conversation-level multi-turn evaluation. The same LLM-as-judge approach AI-Vic runs live against production traffic.
Codex Systems Lab inference fine-tuning python
Research lab for inference-performance benchmarking, light fine-tuning experiments, and agentic system behavior, focused on Codex-style AI coding models. Reproducible experiments for AI coding workflows.
Codex Evaluation Benchmark eval causal-inference python
End-to-end analytics and evaluation pipeline for AI-assisted coding: developer telemetry simulation, LLM code-generation quality scoring, and causal inference on productivity metrics.
OpenAI Foundations openai tutorials
Tutorials and demos for building real-world applications with OpenAI APIs, covering function calling, RAG, embeddings, and multimodal applications through four progressive labs from API basics to production patterns.
LATAM GenAI Lakehouse Benchmark latam spark databricks
Lakehouse-native evaluation framework measuring regional Spanish LLM performance (El Salvador vs Peru) using Delta tables, Spark, and Databricks. Applies Bronze/Silver/Gold data architecture to LLM benchmarking at scale.
Graduate research
UC Berkeley MIDS · 2021 – 2023Projects completed during the UC Berkeley Master of Information and Data Science (MIDS) program.
ML System Engineering & MLOps mlops kubernetes
End-to-end ML platform built on Kubernetes and microservices, including containerized model serving, automated retraining pipelines, CI/CD for ML, and production monitoring. Stack: Kubernetes, Docker, FastAPI, MLflow.
Machine Learning at Scale: Flight Delay Prediction spark databricks
Distributed ML pipeline predicting flight delays across 30M+ records using MapReduce, Hadoop, and Apache Spark on Databricks. Applied ensemble methods (GBT, Random Forest) with feature engineering on temporal and weather data.
Machine Learning: Understanding Hate Crime Patterns tensorflow regression
Applied linear regression and TensorFlow to identify socioeconomic and demographic predictors of hate crime rates across U.S. counties. Surfaced statistically significant correlations to inform policy research.
Data Engineering: Location Recommendations with NoSQL neo4j mongodb redis
Multi-database recommendation engine using Neo4j (graph traversal), MongoDB (document store), and Redis (caching) to generate personalized store location suggestions at low latency.
Data Analysis: NFL Big Data Bowl eda pandas
Exploratory data analysis on NFL tracking data using Python, NumPy, and Pandas. Analyzed player movement patterns and derived game-level insights from raw positional data.
Statistical Analysis: Movie Revenue Regression Study stats ols
Designed and executed a research study on movie revenue predictors using OLS regression, hypothesis testing, and diagnostic analysis to identify drivers of box office performance.
Data Visualization: Travel Guide Reimagined tableau
Interactive Tableau dashboard reimagining travel data as a visual guide, layering geographic, seasonal, and sentiment data to surface non-obvious destination insights.
Capstone: enRoute, Running Route Safety App ios ml
iOS app leveraging real-time safety data and ML-based route scoring to recommend safe running routes. Full mobile + backend stack built as UC Berkeley MIDS capstone project.