Venkata Ramana Reddy
New York, NY · USA · open to remote, hybrid, or relocation
I build production AI that knows your data — not just the internet. Agents, RAG, and the evaluation layer that makes them trustworthy, in enterprise and life-science settings.
- Role
- AI Engineer · agents, RAG, evals
- Experience
- 4 years (2 in production LLM systems)
- Now
- AI Engineer @ FedEx · Co-Founder @ Notes9
- Location
- USA · New York, NY · open to remote, hybrid, or relocation
- Education
- M.S. Artificial Intelligence, University at Buffalo
- Status
- Open to senior AI / LLM engineering roles
The loop I run on every system, from discovery to production
Click a step, or use ← → . Each one is paired with a concrete example from my work so you can see the practice, not just the principle.
Discover
Sit with the people who'll use the system. Find the real workflow, the failure that hurts, and what 'good' would measurably look like.
At FedEx I'm the technical point of contact for a 30-person operations team, each managing 100+ client accounts. Their review sessions decide what the system does next. At Notes9 I lead customer discovery and onboarding with research teams.
Deep builds on real, messy data
End-to-end systems with an evaluation layer and measurable outcomes — from a research agent for life-science teams to auditable refund automation at FedEx.
A production multi-tenant, multi-agent system for life-science teams. Catalyst grounds every answer in a team's own research graph — literature, protocols, experiments, samples, lab notes — instead of the public internet alone.
Research teams keep protocols, experiments, samples, and papers in disconnected tools. General chat assistants answer from the public internet and can't cite a lab's own results, so scientists don't trust them for real work.
Every claim Catalyst makes links to a lab note, experiment, or paper. Retrieval precision and recall are measured by an LLM-as-Judge harness before each release, and user feedback from onboarding feeds directly into agent behavior.
- Built end-to-end on AWS Bedrock and the Anthropic API with a hand-built ReAct tool-use loop.
- Shipped NLP-to-SQL agents over user data and a multi-stage hybrid-ranked RAG pipeline federating live search across PubMed, Europe PMC, and OpenAlex.
- Engineered per-claim, span-level citations (Anthropic Citations plus heuristic grounding) so every claim links to its source.
- LLM-as-Judge eval harness on retrieval precision/recall catches regressions before release.
- Lead customer discovery and onboarding with research teams, turning feedback directly into product and agent behavior changes.
Live literature federation depends on third-party APIs (PubMed, Europe PMC, OpenAlex), so latency varies with their availability. Span-level citation falls back to heuristic grounding when the model returns no native citation.
Where I've worked
- Co-Founder · Notes9Nov 2025 — PresentUnited States
Built Catalyst, a multi-tenant research agent on AWS Bedrock + Anthropic API. Lead customer discovery and onboarding with research teams.
see the case study - AI Engineer · FedExMar 2025 — PresentUnited States
Own the LLM tool-calling system that automates tariff refunds on Azure AI Foundry. Technical point of contact for a 30-person ops team.
see the case study - AI/ML Engineer · FundaeAug 2024 — Feb 2025United States
Delivered a multi-tenant enterprise RAG platform for clients including a top-tier pharma company; built evals, telemetry, and CI/CD.
see the case study - Machine Learning Engineer · Groovy WebJan 2022 — Aug 2023India
End-to-end ML pipelines (TensorFlow, Scikit-learn, XGBoost) with A/B testing and drift monitoring; PySpark/Hive ELT across distributed data lakes.
- AWS Certified Generative AI Developer — ProfessionalAWS · Jul 2026 — Jul 2029
- Building with Claude APIAnthropic · Jul 2026
AI engineer who ships systems teams can trust
AI Engineer with 4 years of experience building reliable, client-facing AI systems in enterprise and life-science settings. I own the full path from requirements with stakeholders to production deployment and ongoing iteration with end users — hands-on across agents, RAG, evals, and AWS/Azure infrastructure.
I care most about reliability, explainability, and the evaluation layer that turns a demo into something a real team depends on — whether that team is a 30-person operations group at FedEx or a research lab using Notes9.