<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>AI Engineering on Manvendra Rajpoot</title>
    <link>https://blog.rajpoot.dev/posts/ai/</link>
    <description>Recent content in AI Engineering on Manvendra Rajpoot</description>
    <image>
      <title>Manvendra Rajpoot</title>
      <url>https://blog.rajpoot.dev/img/personal/cover.png</url>
      <link>https://blog.rajpoot.dev/img/personal/cover.png</link>
    </image>
    <generator>Hugo</generator>
    <language>en-US</language>
    <copyright>Manvendra Rajpoot</copyright>
    <atom:link href="https://blog.rajpoot.dev/posts/ai/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Self-Hosting LLMs in 2026 — When the Math Actually Works</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-self-host-economics-2026/</link>
      <pubDate>Tue, 05 May 2026 08:30:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-self-host-economics-2026/</guid>
      <description>Self-hosting LLMs in 2026 — vLLM, GPU economics, break-even, and when self-host beats API.</description>
    </item>
    <item>
      <title>Anthropic API Best Practices in 2026 — Caching, Tool Use, Streaming, and Production Patterns</title>
      <link>https://blog.rajpoot.dev/posts/ai/anthropic-api-best-practices-2026/</link>
      <pubDate>Tue, 05 May 2026 07:50:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/anthropic-api-best-practices-2026/</guid>
      <description>Anthropic API best practices in 2026 — prompt caching, tool use, streaming, batch API, and production patterns from real Claude apps.</description>
    </item>
    <item>
      <title>Evaluating AI Coding Tools in 2026 — Benchmarks That Matter and Ones That Don&#39;t</title>
      <link>https://blog.rajpoot.dev/posts/ai/ai-coding-evals-2026/</link>
      <pubDate>Tue, 05 May 2026 06:20:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/ai-coding-evals-2026/</guid>
      <description>Evaluating AI coding tools in 2026 — SWE-bench, real-world tasks, and what&amp;#39;s actually predictive of productivity gains.</description>
    </item>
    <item>
      <title>Synthetic Data with LLMs in 2026 — Use Cases, Risks, and the Patterns That Work</title>
      <link>https://blog.rajpoot.dev/posts/ai/synthetic-data-2026/</link>
      <pubDate>Tue, 05 May 2026 06:10:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/synthetic-data-2026/</guid>
      <description>Synthetic data generation with LLMs in 2026 — when it helps, model collapse risk, eval set generation, and production patterns.</description>
    </item>
    <item>
      <title>Voice Agents in 2026 — STT, LLM, TTS, and Latency That Doesn&#39;t Hurt</title>
      <link>https://blog.rajpoot.dev/posts/ai/voice-agents-2026/</link>
      <pubDate>Tue, 05 May 2026 06:00:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/voice-agents-2026/</guid>
      <description>Building voice AI agents in 2026 — streaming STT, LLM, TTS pipelines, latency budgets, interruption, and real-world architectures.</description>
    </item>
    <item>
      <title>Model Context Protocol (MCP) in 2026 — What It Solved, What It Didn&#39;t</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-mcp-protocol-2026/</link>
      <pubDate>Mon, 04 May 2026 06:10:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-mcp-protocol-2026/</guid>
      <description>MCP in 2026 — protocol overview, server / client patterns, ecosystem, and an honest take on where MCP fits in agent infrastructure.</description>
    </item>
    <item>
      <title>LLM Tool Use Patterns in 2026 — Schemas, Validation, and the Loop</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-tool-use-patterns-2026/</link>
      <pubDate>Mon, 04 May 2026 06:00:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-tool-use-patterns-2026/</guid>
      <description>LLM tool use in 2026 — designing tool schemas, parallel calls, error handling, and the patterns from production agents.</description>
    </item>
    <item>
      <title>Agentic Coding in 2026 — Claude Code, Cursor, and the Real Workflow</title>
      <link>https://blog.rajpoot.dev/posts/ai/agentic-coding-2026/</link>
      <pubDate>Sun, 03 May 2026 08:50:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/agentic-coding-2026/</guid>
      <description>Agentic coding in 2026 — Claude Code, Cursor, Aider, and how AI coding agents actually fit into senior engineers&amp;#39; workflows.</description>
    </item>
    <item>
      <title>LLM Batch Processing in 2026 — Anthropic / OpenAI Batch API for 50% Off</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-batch-processing-2026/</link>
      <pubDate>Sun, 03 May 2026 06:20:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-batch-processing-2026/</guid>
      <description>LLM batch APIs in 2026 — Anthropic, OpenAI, Bedrock batch processing for 50% discount, when to use them, and the patterns that work.</description>
    </item>
    <item>
      <title>LLM Deployment Patterns in 2026 — Inference Servers, Routing, and Production Architectures</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-deployment-patterns-2026/</link>
      <pubDate>Sun, 03 May 2026 06:10:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-deployment-patterns-2026/</guid>
      <description>LLM deployment patterns in 2026 — vLLM, TGI, Ollama, hybrid API&#43;self-hosted, routing layers, and the production architectures that actually work.</description>
    </item>
    <item>
      <title>Prompt Engineering in 2026 — What Still Works, What Doesn&#39;t, and What Changed</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-prompt-engineering-2026/</link>
      <pubDate>Sun, 03 May 2026 06:00:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-prompt-engineering-2026/</guid>
      <description>Prompt engineering in 2026 — patterns that still work, what&amp;#39;s been obsoleted by better models, structured prompts, and production discipline.</description>
    </item>
    <item>
      <title>LLM Agent Frameworks in 2026 — LangGraph, CrewAI, and the Bare-Metal Alternative</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-agent-frameworks-2026/</link>
      <pubDate>Sat, 02 May 2026 12:50:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-agent-frameworks-2026/</guid>
      <description>LLM agent frameworks in 2026 — LangGraph, CrewAI, OpenAI Agents SDK, AutoGen, and when bare-metal is better.</description>
    </item>
    <item>
      <title>Agent Memory Systems in 2026 — Episodic, Semantic, and the Patterns That Stick</title>
      <link>https://blog.rajpoot.dev/posts/ai/agent-memory-systems-2026/</link>
      <pubDate>Sat, 02 May 2026 11:10:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/agent-memory-systems-2026/</guid>
      <description>Agent memory systems in 2026 — episodic vs semantic memory, vector stores, working memory, and patterns from production agents.</description>
    </item>
    <item>
      <title>LLM Context Windows in 2026 — Long Context, Cache, and the Limits of &#39;Just Add More&#39;</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-context-windows-2026/</link>
      <pubDate>Sat, 02 May 2026 11:00:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-context-windows-2026/</guid>
      <description>LLM context windows in 2026 — what 200k / 1M context can and can&amp;#39;t do, prompt caching, retrieval, and patterns from production.</description>
    </item>
    <item>
      <title>Multimodal LLMs in 2026 — Vision, Audio, and What&#39;s Actually Useful</title>
      <link>https://blog.rajpoot.dev/posts/ai/multimodal-llms-2026/</link>
      <pubDate>Sat, 02 May 2026 09:30:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/multimodal-llms-2026/</guid>
      <description>Multimodal LLMs in 2026 — vision input, audio input, generation, real-world use cases, and the patterns that work in production.</description>
    </item>
    <item>
      <title>Evaluating RAG Systems in 2026 — Retrieval Quality, Faithfulness, and the Metrics That Matter</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-rag-evaluation-2026/</link>
      <pubDate>Sat, 02 May 2026 09:20:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-rag-evaluation-2026/</guid>
      <description>RAG evaluation in 2026 — retrieval metrics (recall, MRR), generation metrics (faithfulness, relevance), Ragas, and the patterns from production RAG.</description>
    </item>
    <item>
      <title>LLM Observability in 2026 — Tracing, Evals, and the Things You Can&#39;t Skip</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-observability-2026/</link>
      <pubDate>Sat, 02 May 2026 07:30:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-observability-2026/</guid>
      <description>Production LLM observability in 2026 — distributed tracing, eval pipelines, Langfuse, Arize, and the patterns that turn black-box LLMs into operable systems.</description>
    </item>
    <item>
      <title>LLM Cost Optimization in 2026 — From Bills That Hurt to Bills That Don&#39;t</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-cost-optimization-2026/</link>
      <pubDate>Sat, 02 May 2026 07:20:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-cost-optimization-2026/</guid>
      <description>Cutting LLM costs in 2026 — prompt caching, routing, batching, fine-tunes, and the patterns that drop bills 5-20× without quality loss.</description>
    </item>
    <item>
      <title>LLM Guardrails in 2026 — Input Filtering, Output Validation, and Safety Nets</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-guardrails-content-safety-2026/</link>
      <pubDate>Fri, 01 May 2026 07:10:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-guardrails-content-safety-2026/</guid>
      <description>Practical LLM guardrails in 2026 — input filtering, output validation, NVIDIA NeMo, Guardrails AI, and the patterns that prevent embarrassments.</description>
    </item>
    <item>
      <title>Embedding Databases in 2026 — pgvector, Qdrant, Weaviate, Milvus, Pinecone</title>
      <link>https://blog.rajpoot.dev/posts/ai/embedding-databases-2026/</link>
      <pubDate>Fri, 01 May 2026 07:00:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/embedding-databases-2026/</guid>
      <description>Embedding databases compared in 2026 — pgvector, Qdrant, Weaviate, Milvus, Pinecone, Vectorize. When each fits.</description>
    </item>
    <item>
      <title>Fine-Tuning LLMs in 2026 — LoRA, QLoRA, and the Cheap Path to Specialized Models</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-fine-tuning-lora-qlora-2026/</link>
      <pubDate>Fri, 01 May 2026 06:00:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-fine-tuning-lora-qlora-2026/</guid>
      <description>Practical LLM fine-tuning in 2026 — LoRA, QLoRA, training data prep, evaluation, and the patterns from teams shipping fine-tuned models.</description>
    </item>
    <item>
      <title>LLM Agent Error Recovery in 2026 — Patterns That Don&#39;t Loop Forever</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-agent-error-recovery-2026/</link>
      <pubDate>Fri, 01 May 2026 04:30:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-agent-error-recovery-2026/</guid>
      <description>How to build LLM agents that recover from errors gracefully — retry policies, fallback paths, max-step caps, and the patterns that prevent runaway loops.</description>
    </item>
    <item>
      <title>OpenAI vs Anthropic vs Google for Production AI in 2026</title>
      <link>https://blog.rajpoot.dev/posts/ai/openai-vs-anthropic-vs-google-2026/</link>
      <pubDate>Fri, 01 May 2026 03:00:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/openai-vs-anthropic-vs-google-2026/</guid>
      <description>Honest comparison of OpenAI vs Anthropic vs Google for production LLM apps in 2026 — model quality, pricing, latency, ecosystem, and how to pick.</description>
    </item>
    <item>
      <title>Document AI in 2026 — Extracting Structured Data from PDFs and Images</title>
      <link>https://blog.rajpoot.dev/posts/ai/document-ai-pdf-extraction-2026/</link>
      <pubDate>Fri, 01 May 2026 01:50:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/document-ai-pdf-extraction-2026/</guid>
      <description>How to extract structured data from PDFs and images in 2026 — vision LLMs, OCR pipelines, layout-aware models, and the patterns that ship.</description>
    </item>
    <item>
      <title>LLM Prompt Caching Deep Dive — Anthropic, OpenAI, and the Patterns That Save 90%</title>
      <link>https://blog.rajpoot.dev/posts/ai/llm-prompt-caching-deep-dive-2026/</link>
      <pubDate>Fri, 01 May 2026 00:30:00 +0530</pubDate>
      <guid>https://blog.rajpoot.dev/posts/ai/llm-prompt-caching-deep-dive-2026/</guid>
      <description>Prompt caching mechanics in 2026 — Anthropic&amp;#39;s ephemeral cache, OpenAI&amp;#39;s automatic caching, breakpoint placement, hit-rate measurement, and the patterns that save real money.</description>
    </item>
  </channel>
</rss>
