Victor Ramirez · San Francisco Bay Area · Director, Developer & Platform Experience @ Moody's Analytics
I build production AI systems and the platforms that let hundreds of engineers ship them.
RAG pipelines, LLM evaluation frameworks, and agentic workflows: measured, shipped, and operated in production. UC Berkeley MIDS. I teach the same material hands-on through Techqueria and the AI Builders: LatinX Edition podcast.
- 17 yrs
- production systems at scale
- 3 yrs
- AI architecture & evaluation
- 3
- live demos you can use right now
- 6+
- podcast episodes & 3 conference talks
- 0
- fine-tunes shipped — I trained a 7B LoRA adapter on AI-Vic's eval data, A/B tested it against my own judge, and the base model won
What I build
shipped to productionRAG pipelines & retrieval systems
Retrieval-augmented generation connecting LLMs to enterprise knowledge at scale.
LLM evaluation frameworks
Recall@k, Precision@k, MRR, groundedness. Rigorous benchmarking with measurable metrics.
Agentic workflows
Multi-turn orchestration, function calling, and MCP tool-use patterns in production.
Developer platforms
Internal platforms, observability rollouts, and incident coordination that keep engineering teams shipping and systems reliable.
Featured work
github.com/ramirez-ai-labsLatino Canon live rag eval mcp
Bilingual hybrid search over 294 Latino-led films and series. A retrieval eval checks every deploy (recall@5 0.818), an LLM judge gates every generated note, and a remote MCP server works in Claude. $0/month on Cloudflare.
TrustClaw live claude vercel
Autonomous email-summarization agent: Gmail webhook + Claude, every dependency proxied through JFrog Artifactory, with a supply-chain audit trail across 900+ packages.
AI Operating System live agents eval langraph
Four-domain LangGraph orchestration with Claude tool-loop safety circuit breaker, CI-gated eval harness, and deterministic fallback paths. Production-grade agentic system grounding.
RAG Evaluation Lab rag eval python
Fully offline lab for evaluating RAG systems: synthetic datasets, embeddings, vector search, and the retrieval metrics I teach at Techqueria.
Speaking
2025 – 2026Writing
medium.com/@vhr1975I Fine-Tuned a Model for My Portfolio Chatbot. My Own Eval Told Me Not to Ship It. new ai-vic llm-eval fine-tuning
AI-Vic runs its own LLM-as-judge eval in production: a small model scores every reply for relevance and groundedness, nightly and on every visitor 👍/👎. I used that eval's output to fine-tune a 7B LoRA adapter, A/B tested every version against the un-adapted base with the same judge, and it regressed on every run — so I shipped the base model. The write-up covers the vacuous-5/5 bug the eval had first, the hard-case rubrics that fixed it, and why the negative result is the point.
The Math We All Half-Remember Is Quietly Running Every LLM You Use new embeddings fundamentals
Why sine and cosine show up inside GPT, Claude, and LLaMA: connecting cosine similarity — the metric every embeddings workshop leans on — back to the spinning unit circle it comes from. AI as less magic, more math you already half-know.
What Dario Amodei's Two AI Essays Taught Me About Building AI Platforms Inside a Company ai-platform governance
Translating Amodei's essays on AI potential and risk down from frontier-research scale to the one I actually work at: getting enterprise engineering teams to adopt AI safely, one platform onboarding at a time.
RAG Evaluation, Part 1: Fundamentals Without the Cloud rag llm-eval
Part 1 of a 3-part series. The offline baseline: keyword and TF-IDF retrieval measured with Recall@k, Precision@k, and MRR, no cloud infrastructure. Honest numbers to build on.
RAG Evaluation, Part 2: Does Embedding Quality Justify the Cost? rag llm-eval vertex-ai
A Vertex AI benchmark with real production numbers. TF-IDF gets Recall@3 of 0.833 for free in under 5ms; embeddings hit 1.0 but cost money and run 50× slower. What the trade-off actually looks like.
RAG Evaluation, Part 3: The Retriever Decision Framework rag llm-eval
Turning a customer constraint — timeline, budget, compliance — into a retriever recommendation you can defend, instead of picking off a benchmark leaderboard.
What I learned Building a Perplexity API api llm integration
Building with the Perplexity API: lessons on real-time search integration, response streaming, and practical patterns for production LLM applications.
Meet AI-Vic: A Conversational Version of My Portfolio ai-vic llm product
The original launch post for this site's portfolio AI assistant: system prompt grounding, edge inference with Llama 3.3, and conversational UX design. (AI-Vic has since grown a full RAG, agentic tool-use, and LLM-as-judge eval layer — see the follow-ups above.)
Demystifying Generative AI: My Journey with Techqueria Workshops teaching genai
Behind the curriculum: how I designed hands-on AI workshops for the LatinX engineering community, from LLM fundamentals to production RAG systems and agentic architectures.
Introducing AI Builders: LatinX Edition community podcast
The story behind the podcast: why I started a show amplifying LatinX voices in AI and what I've learned from conversations with practitioners building at the frontier of generative AI.