<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Refactor — Notes from the AI engineering frontier</title>
    <link>https://refactor4ai.com/blog</link>
    <atom:link href="https://refactor4ai.com/blog/rss.xml" rel="self" type="application/rss+xml"/>
    <description>Practical writing on AI fluency for senior engineers, platform leads, and product managers.</description>
    <language>en-gb</language>
    <lastBuildDate>Sun, 28 Jun 2026 11:45:22 GMT</lastBuildDate>
    <item>
      <title>Why We Don&apos;t Use Agent Frameworks — and What We Built Instead</title>
      <link>https://refactor4ai.com/blog/why-we-dont-use-agent-frameworks</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/why-we-dont-use-agent-frameworks</guid>
      <pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate>
      <description>We run agents in production on the raw provider SDK plus about 300 lines of code we own — not LangGraph, CrewAI, the OpenAI Agents SDK, or Google ADK. The frameworks are genuinely good for prototyping, but in production their orchestration abstractions hide the three things you most need to see: the exact prompt, the exact control flow, and the exact tool calls. Here&apos;s the honest case for and against, the loop we actually ship, and when you should pick a framework anyway.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Versioning Prompts Like Code</title>
      <link>https://refactor4ai.com/blog/versioning-prompts-like-code</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/versioning-prompts-like-code</guid>
      <pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate>
      <description>Your prompts are production logic, and right now they&apos;re living in Slack threads, inlined string literals, and someone&apos;s local notes. Treat them as versioned artifacts: every change tracked, label-based deployment that decouples editing from releasing, a one-line reason captured at write time, and an eval gate in CI. Here&apos;s the discipline and the tool decision — raw git for solo devs, a workbench like Langfuse or PromptLayer when non-engineers edit prompts.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>SLM vs LLM: When to Ship a 7B Model</title>
      <link>https://refactor4ai.com/blog/slm-vs-llm-when-to-ship-7b</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/slm-vs-llm-when-to-ship-7b</guid>
      <pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate>
      <description>Most teams reach for a frontier model out of habit when a 7–8B small language model would be cheaper, faster, private, and good enough for the task in front of them. Phi-4, Gemma 3, Qwen2.5-Coder, and Llama 3.1 8B clear real production bars on narrow, structured work. Here&apos;s what small models are genuinely good at, where they fall down, and the decision rule for when to ship a 7B instead of paying frontier prices.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The AI Risk Register Every PM Should Keep</title>
      <link>https://refactor4ai.com/blog/ai-risk-register-pm</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-risk-register-pm</guid>
      <pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate>
      <description>If you ship AI features and can&apos;t produce a risk register on demand, you&apos;re one incident — or one EU AI Act audit — away from a very bad week. Build it to NIST AI RMF from day one, map each entry to its EU AI Act article, and one living document becomes your evidence base for every framework at once. Here are the ten risks that belong on it, the columns that make it auditable, and how to actually run it.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AI Feature ROI: The Math You&apos;ll Be Asked</title>
      <link>https://refactor4ai.com/blog/ai-feature-roi-math</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-feature-roi-math</guid>
      <pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate>
      <description>When an exec asks &apos;is it worth it?&apos;, the answer is a number, and the part that surprises everyone is the cost side — agentic features burn 5–30× the tokens of a chatbot, so the request-cost assumption you carried over is wrong by an order of magnitude. Here&apos;s the ROI formula, the cost model that doesn&apos;t blow up in production, a fully worked example, and the three inputs that move the result more than which model you pick.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>PII Scrubbing in the Prompt Path</title>
      <link>https://refactor4ai.com/blog/pii-scrubbing-in-prompt-path</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/pii-scrubbing-in-prompt-path</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>Scrub PII before the prompt leaves your network — and scrub the retrieved chunks too, which is the step almost everyone skips. A production pattern built on Microsoft Presidio: where the leaks actually happen, reversible vs irreversible redaction, where to place the proxy (LiteLLM, PII Shield), and the chunk-redaction gap that quietly ships customer data to vendor logs.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Picking a Cloud for Your First AI Product</title>
      <link>https://refactor4ai.com/blog/picking-a-cloud-for-first-ai-product</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/picking-a-cloud-for-first-ai-product</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>If you&apos;re already on a cloud, build your first AI product there — the IAM, VPC, and data-gravity integration compounds, and the model catalog won&apos;t lock you in. If you&apos;re greenfield, let the one model you can&apos;t live without decide: Bedrock for the broadest catalog, Vertex for Gemini-native and startup cost, Azure Foundry for day-one OpenAI and the Microsoft stack. The 2026 comparison, and the thing that actually locks you in.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The Economics of Agents: Cost-Per-Task Math</title>
      <link>https://refactor4ai.com/blog/economics-of-agents-cost-per-task</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/economics-of-agents-cost-per-task</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>Budget agents per completed task, not per token — an agentic task burns 5–30× the tokens of a single chatbot call, and the multi-turn loop makes it worse than linear. Here&apos;s the cost-per-task formula, where the tokens actually go, a worked example that lands a coding agent at ~$3.40 a task, and the four levers that cut it by 60–80% without touching answer quality.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AI in Legacy Code: How Teams Are Actually Modernizing</title>
      <link>https://refactor4ai.com/blog/ai-in-legacy-code-modernization</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-in-legacy-code-modernization</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>Point AI at comprehension first, not code generation — on a legacy codebase the bottleneck is understanding what the system does, not typing the replacement. The four jobs LLMs genuinely do well on old code, what IBM watsonx Code Assistant for Z is doing to COBOL, why characterization tests are the unlock, and the staged plan that beats the big-bang rewrite that has killed modernization programs for thirty years.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Senior vs Staff vs Principal AI Engineer — What Each Ladder Looks For</title>
      <link>https://refactor4ai.com/blog/ai-engineer-ladders-explained</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-engineer-ladders-explained</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>What gets you promoted on the AI engineering ladder is the blast radius of the decisions you&apos;re trusted to make without a net — not how much you know about transformers. Senior owns a system, Staff owns a problem across teams, Principal owns a direction for the org. The scope, the comp bands (US and AI-startup, 2026), the AI-specific bar, and the one promotion mistake at each level.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Semantic Caching: The Real Numbers on Latency and Cost</title>
      <link>https://refactor4ai.com/blog/semantic-caching-real-numbers</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/semantic-caching-real-numbers</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>A semantic cache in front of your LLM serves the answer to a question you&apos;ve already answered — even when it&apos;s phrased differently — for the price of one embedding lookup. On repetitive traffic that removes 30–60% of calls and collapses latency from seconds to milliseconds. The whole engineering problem is the similarity threshold: set it at 0.92, not 0.80, or you&apos;ll confidently return the wrong cached answer.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Self-Hosted vs Bedrock vs Azure OpenAI: The TCO Model (2026)</title>
      <link>https://refactor4ai.com/blog/self-hosted-vs-bedrock-vs-azure-tco</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/self-hosted-vs-bedrock-vs-azure-tco</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>The cloud-marketplace markup for routing a model through Bedrock or Azure OpenAI instead of the vendor&apos;s direct API is only 10–20% — and it buys VPC peering, unified billing, and compliance you&apos;d otherwise build yourself. Self-hosting looks cheaper on a GPU spec sheet and is 3–5x more expensive all-in until you&apos;re very large. Here&apos;s the three-column TCO model that makes the call.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>How We Built Refactor4AI with Claude</title>
      <link>https://refactor4ai.com/blog/how-we-built-refactor4ai-with-claude</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/how-we-built-refactor4ai-with-claude</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>We built this site — content engine included — mostly with Claude, and the decision that mattered wasn&apos;t the code generation. It was treating the content pipeline as a set of scheduled agents gated by an eval pass, not a CMS. Posts are React components, the backlog is a spreadsheet, and a nightly agent drafts while a separate review agent scores and routes to publish. Here&apos;s the architecture and what we&apos;d keep, change, and never automate.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Bedrock Agents vs Azure AI Agents vs Vertex AI Agent Builder (2026)</title>
      <link>https://refactor4ai.com/blog/bedrock-vs-azure-vs-vertex-agents</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/bedrock-vs-azure-vs-vertex-agents</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>Each hyperscaler now ships its own managed agent runtime, and the demos look identical. The decision isn&apos;t about the demo — it&apos;s about where your identity, data, and on-call already live. Bedrock AgentCore for AWS-native shops, Azure AI Foundry Agent Service for anyone deep in Microsoft 365, Vertex AI Agent Builder when you want control over the loop and Google Search grounding. Here&apos;s the real comparison.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AI Startup Unit Economics: The Model to Build Before You Raise</title>
      <link>https://refactor4ai.com/blog/ai-startup-unit-economics</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-startup-unit-economics</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>AI startups don&apos;t have SaaS margins, and pretending otherwise is how you raise on a number you can&apos;t hit. The 2026 benchmark for AI product gross margin is ~52% against the old 80% SaaS bar, inference eats ~23% of revenue at scaling-stage companies, and your unit of cost is the inference, not the seat. Here&apos;s the per-user model investors actually underwrite, and the four levers that move it.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Pricing AI Features in 2026: Per-Seat vs Usage vs Hybrid (and Why Everyone Just Switched)</title>
      <link>https://refactor4ai.com/blog/pricing-ai-features-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/pricing-ai-features-2026</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>If your AI feature runs agents, flat per-seat pricing is quietly bankrupting you on your heaviest users. In one 72-hour stretch in June 2026, Copilot, Cursor, and Claude Code all moved to metered billing for the same reason. Here&apos;s the per-seat vs usage vs hybrid trade-off, the margin math to run first, and the packaging that actually holds up.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Multimodal RAG in 2026: When Text-Only Retrieval Quietly Fails Your PDFs</title>
      <link>https://refactor4ai.com/blog/multimodal-rag-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/multimodal-rag-2026</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>Roughly 80% of enterprise PDFs contain a table, chart, or complex layout — and a text-only RAG pipeline drops all of it on the floor. On financial documents, visual retrieval lifts recall from ~62% to ~84%. Here are the three multimodal RAG architectures that matter in 2026, the storage trade-off nobody budgets for, and what we&apos;d actually ship.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Designing Trust Signals for AI Features: Make It Verifiable, Not Confident</title>
      <link>https://refactor4ai.com/blog/designing-trust-signals-ai</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/designing-trust-signals-ai</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>Users distrust AI output for a good reason — it&apos;s confidently wrong often enough to have earned the skepticism. The fix isn&apos;t a more assured tone or decorative citations. It&apos;s designing for verification: load-bearing citations, calibrated confidence, visible provenance, and a one-click path to the source. Here are the patterns that build trust and the ones that fake it.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Code Review When Half the Diff Is Machine-Written</title>
      <link>https://refactor4ai.com/blog/code-review-ai-generated-code</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/code-review-ai-generated-code</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>AI now authors roughly 27% of measured production commits and 40%+ by developer self-report — and your review process was built for human-paced, human-shaped diffs. Here&apos;s how to restructure review for a world of large, plausible, machine-written PRs: a two-layer model with AI reviewers for breadth and humans for the judgment no model can outsource.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Cheap-First Cascade Routing: Try the $0.40 Model Before You Pay for the $30 One</title>
      <link>https://refactor4ai.com/blog/cheap-first-cascade-routing</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/cheap-first-cascade-routing</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>Most AI features overpay because they send every request to a frontier model. A cheap-first cascade — start on a speed-tier model, escalate only what fails a confidence gate — cuts 60–80% of inference spend at the same answer quality. Here&apos;s the ladder, the escalation logic that actually works, and whether to build it or buy a gateway.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>TPUs vs H100s vs Trainium (2026): Who Wins What</title>
      <link>https://refactor4ai.com/blog/tpus-vs-h100-vs-trainium-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/tpus-vs-h100-vs-trainium-2026</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>The chip with the best spec sheet rarely wins. NVIDIA Blackwell is the default because you can rent it anywhere and CUDA runs everything; Google&apos;s Ironwood TPU wins price-per-token if you live in JAX and GCP; AWS Trainium is an inference-cost lever inside the AWS walled garden. A practical 2026 buyer&apos;s guide for the chip decision that&apos;s really driving your AI bill.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Stop Using LangChain in Production</title>
      <link>https://refactor4ai.com/blog/stop-using-langchain-production</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/stop-using-langchain-production</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Prototype on LangChain if it helps, but ship production on the raw provider SDK or a thin layer you own. The abstraction tax — opaque control flow, hidden prompts, relentless version churn — now exceeds the value, because the model vendors absorbed the primitives LangChain was invented to provide. Here&apos;s the honest case for leaving, the case for staying, and a migration path.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Prompt Injection: A 2026 Defense Playbook</title>
      <link>https://refactor4ai.com/blog/prompt-injection-defense-playbook-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/prompt-injection-defense-playbook-2026</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>You cannot filter your way to safety. Prompt injection remains unsolved at the model layer, so the only durable defense is architectural: assume the model will be hijacked and make that harmless with least-privilege tools, action gating, and a quarantined-untrusted-content pattern. A concrete, layered playbook mapped to the 2026 OWASP LLM and agentic risk lists.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Building an Eval Harness From Scratch</title>
      <link>https://refactor4ai.com/blog/building-eval-harness-from-scratch</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/building-eval-harness-from-scratch</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>You can stand up a useful LLM eval harness in an afternoon with a CSV and a 40-line script — no platform required. The hard part was never the framework; it&apos;s the golden dataset and wiring evals into CI so a quality regression fails the build instead of shipping silently. Here&apos;s the minimal version that actually moves the needle, and how to grow it.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AI for SQL: Where It Shines, Where It Embarrasses You</title>
      <link>https://refactor4ai.com/blog/ai-for-sql-shines-embarrasses</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-for-sql-shines-embarrasses</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Text-to-SQL is production-ready for exactly one job — assisting an analyst who can read the query before running it — and a liability everywhere else. The 2026 benchmark numbers explain why: on clean schemas top systems clear ~82%, but on real enterprise databases execution accuracy collapses toward 39%. Here&apos;s where to ship it, where to fence it off, and the architecture that makes it safe.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Self-Hosting vs Managed LLM: The Break-Even Math (2026)</title>
      <link>https://refactor4ai.com/blog/self-host-vs-managed-llm-breakeven</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/self-host-vs-managed-llm-breakeven</guid>
      <pubDate>Sat, 20 Jun 2026 00:00:00 GMT</pubDate>
      <description>Most teams that self-host an LLM lose money doing it. The honest break-even against frontier APIs is around 100–256M tokens/month; against open-model API providers it&apos;s 50M+ tokens/day. Here&apos;s the real math, the 3–5x hidden-cost multiplier nobody budgets for, and the three reasons self-hosting is still the right call.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Memory in Agents: Episodic, Semantic, Scratchpad (2026)</title>
      <link>https://refactor4ai.com/blog/memory-in-agents-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/memory-in-agents-2026</guid>
      <pubDate>Sat, 20 Jun 2026 00:00:00 GMT</pubDate>
      <description>Your agent doesn&apos;t have a memory problem — it has four, and conflating them is why it forgets. The 2026 playbook for agent memory: working memory (the scratchpad), episodic (what happened), semantic (what&apos;s true), and procedural (how to act). What each is for, how to store and retrieve it, and the failure mode of dumping everything into one vector store.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>GPU Inference Economics 2026: A100 vs H100 vs H200 vs B200</title>
      <link>https://refactor4ai.com/blog/gpu-inference-economics-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/gpu-inference-economics-2026</guid>
      <pubDate>Sat, 20 Jun 2026 00:00:00 GMT</pubDate>
      <description>The cheapest GPU per hour is almost never the cheapest GPU per token. A field guide to A100, H100, H200, and B200 inference economics in 2026 — real rental prices, why B200 wins on throughput-adjusted cost, the memory math that makes H200 the quiet workhorse, and the utilization trap that wrecks every spreadsheet.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AI Velocity: What &apos;Shipping Fast&apos; Actually Means With LLMs</title>
      <link>https://refactor4ai.com/blog/ai-velocity-shipping-fast</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-velocity-shipping-fast</guid>
      <pubDate>Sat, 20 Jun 2026 00:00:00 GMT</pubDate>
      <description>AI features feel slower to ship not because the code is harder but because the old definition of velocity stops working. With non-deterministic models, velocity isn&apos;t commits per week — it&apos;s confident iterations per week, and the only thing that buys it is an eval harness. Why eval is the velocity unlock, not the tax, and how to stop your best engineers from QA-ing prompts by hand.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The AI Feature Launch Checklist (2026)</title>
      <link>https://refactor4ai.com/blog/ai-feature-launch-checklist</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-feature-launch-checklist</guid>
      <pubDate>Sat, 20 Jun 2026 00:00:00 GMT</pubDate>
      <description>The thing that delays AI launches isn&apos;t the model — it&apos;s the absence of an eval gate, a rollback path, and a guardrail story you can defend. A pre-launch checklist for AI PMs: the seven gates every AI feature should clear before it ships, the regulatory deadlines that are now real, and the one artifact that turns a scary launch into a boring one.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Vector DB Showdown 2026: Pinecone vs Weaviate vs Qdrant vs pgvector</title>
      <link>https://refactor4ai.com/blog/vector-database-showdown-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/vector-database-showdown-2026</guid>
      <pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate>
      <description>Four credible choices, four different reasons to pick them. pgvector if you&apos;re under 10M vectors and already on Postgres, Weaviate for the best hybrid search, Qdrant for price-performance and latency at scale, Pinecone if zero-ops managed is worth the markup. Real 2026 pricing, real benchmark latencies, and the cost trap that puts most teams 2.5–4x over budget.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>RAG Is Dead. Long Live Context Engineering.</title>
      <link>https://refactor4ai.com/blog/rag-is-dead-context-engineering</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/rag-is-dead-context-engineering</guid>
      <pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate>
      <description>Naive RAG — embed everything, top-k cosine, stuff the prompt — is dead, and it deserved to die. But the &apos;just use a 1M-token window&apos; crowd is wrong too. The actual 2026 skill is context engineering: deciding which tokens reach the model at which step. Retrieval is one tool inside that discipline, not the architecture itself.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Function Calling Deep Dive: Schema Design LLMs Love</title>
      <link>https://refactor4ai.com/blog/function-calling-schema-design</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/function-calling-schema-design</guid>
      <pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate>
      <description>Tool-use reliability is decided by your schema, not your model. Strict types, enums over free strings, one tool one job, descriptions written as prompts, and ask-don&apos;t-guess on anything that touches money or deletion. Plus the rule everyone skips: provider-side structured outputs are not validation — treat every tool call as untrusted input.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Chunking Strategies That Don&apos;t Suck (2026)</title>
      <link>https://refactor4ai.com/blog/chunking-strategies-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/chunking-strategies-2026</guid>
      <pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate>
      <description>Start with recursive 512–1024 token chunks and 10–15% overlap — it beats most clever schemes and it&apos;s where the benchmarks land. Skip semantic chunking; it&apos;s ~14x slower for marginal gains. The two upgrades that actually move retrieval are Anthropic&apos;s contextual retrieval (−49% failed retrievals, −67% with reranking) and Jina&apos;s late chunking. Here&apos;s the ladder, in order of ROI.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Azure AI Search as RAG Backbone: When It Beats DIY</title>
      <link>https://refactor4ai.com/blog/azure-ai-search-rag-backbone</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/azure-ai-search-rag-backbone</guid>
      <pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate>
      <description>If you&apos;re already on Azure, default to Azure AI Search for production RAG and reach for a DIY stack only when you have a concrete reason. Here&apos;s the 2026 feature set — integrated vectorization, the semantic ranker, and agentic retrieval — what each meter actually costs you, and the three situations where rolling your own pgvector pipeline still wins.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Long-context vs RAG: when 2M tokens actually beats retrieval</title>
      <link>https://refactor4ai.com/blog/long-context-vs-rag-2m-tokens</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/long-context-vs-rag-2m-tokens</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>Gemini&apos;s 2M-token window tempts teams to delete their RAG pipeline. Sometimes that&apos;s right. Usually it isn&apos;t. Here&apos;s the decision rule — based on corpus size, update frequency, citation needs, and the gap between advertised and effective context — plus the hybrid pattern most teams should run.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Datadog vs Langfuse vs Helicone vs Arize: A Real Comparison</title>
      <link>https://refactor4ai.com/blog/llm-observability-tools-comparison</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/llm-observability-tools-comparison</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>Four LLM observability tools that solve different problems. Helicone is the fastest install (proxy, no SDK changes). Langfuse is the open-source framework-agnostic default. Arize Phoenix scales with enterprise ML telemetry. Datadog keeps LLM spans next to your existing APM. Here&apos;s which to pick and the cost traps to avoid.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The empty-state problem in AI products (and how to design it away)</title>
      <link>https://refactor4ai.com/blog/empty-state-ai-products</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/empty-state-ai-products</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>The blank screen is where most AI products lose the user. A chat box with a blinking cursor is not an onboarding flow. Here&apos;s how to design the first-run experience so users hit value before they have to invent a prompt — seeded content, segmented starts, and in-context nudges.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Embedding Model Selection: Which to Use in 2026</title>
      <link>https://refactor4ai.com/blog/embedding-model-selection-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/embedding-model-selection-2026</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>OpenAI&apos;s text-embedding-3 is no longer the automatic default. In 2026, Voyage leads domain-specific retrieval, Cohere embed-v4 leads multilingual, Gemini Embedding is the cheapest credible option, and BGE-M3 wins self-hosted. Here&apos;s how to pick by use case instead of by brand — and why you should benchmark on your own data.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AWS Bedrock Knowledge Bases: Production-Grade Managed RAG?</title>
      <link>https://refactor4ai.com/blog/bedrock-knowledge-bases-production-rag</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/bedrock-knowledge-bases-production-rag</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>Bedrock Knowledge Bases will get a working RAG pipeline live in an afternoon and handle ingestion, chunking, embeddings, and retrieval for you. The catch is a ~$700/month OpenSearch Serverless floor and a ceiling on retrieval tuning. Here&apos;s when managed RAG is the right call and when to roll your own.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AI moats in 2026: what&apos;s defensible, what&apos;s just a feature</title>
      <link>https://refactor4ai.com/blog/ai-moats-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-moats-2026</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>Model access isn&apos;t a moat — it&apos;s a commodity that reprices every quarter. The defensible things in 2026 are proprietary data flywheels, deep workflow integration, regulatory barriers, and operational reliability nobody can clone in a weekend. Here&apos;s the honest map of what defends an AI product and what gets copied.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The AI Metrics Cheatsheet for PMs: RAG, Agents, and Generative</title>
      <link>https://refactor4ai.com/blog/ai-metrics-cheatsheet-pm</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-metrics-cheatsheet-pm</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>Most AI dashboards measure the wrong things. This is the metrics cheatsheet we hand product managers — what to track for RAG, for agents, and for generative features, which numbers actually predict user trust, and the three KPIs that belong on every exec slide.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The AI Engineering Roadmap for 2026 (No PhD Required)</title>
      <link>https://refactor4ai.com/blog/ai-engineering-roadmap-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-engineering-roadmap-2026</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>A realistic 8–12 month path from competent software engineer to employable AI engineer in 2026. Five ordered stages — Python and LLM APIs, prompt engineering, production RAG, agents, and deployment — with the checkpoints that actually prove you can ship. A portfolio of deployed projects beats a certificate.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AI Engineer Salary Bands 2026: What the Market Actually Pays</title>
      <link>https://refactor4ai.com/blog/ai-engineer-salary-bands-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-engineer-salary-bands-2026</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>The 2026 AI engineering market has split into two economies — enterprise comp of $170K–$245K total, and a frontier-lab cohort clearing $600K–$1M+ for the same job title. Here are the real bands by level, the specialties that move the number, and how to read an offer.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Agent golden paths: platform-team patterns for shipping AI safely</title>
      <link>https://refactor4ai.com/blog/agent-golden-paths-platform</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/agent-golden-paths-platform</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>If your platform team treats AI agents as a special case, you&apos;ll drown in shadow tooling and one-off approvals. Treat an agent as just another user persona with RBAC, quotas, and a paved road. Here are the golden-path patterns that let agents ship without raising your blast radius.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>What we wish we knew before our first agent shipped</title>
      <link>https://refactor4ai.com/blog/wish-we-knew-before-first-agent</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/wish-we-knew-before-first-agent</guid>
      <pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate>
      <description>Nine specific things — pre-deployment evals lie, tool latency dominates token cost, the model will confidently fabricate from a half-broken tool call — that we learned in production and would now tell every team about to ship their first agent.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Voice agents in 2026: the sub-500ms stack</title>
      <link>https://refactor4ai.com/blog/voice-agents-sub-500ms-stack</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/voice-agents-sub-500ms-stack</guid>
      <pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate>
      <description>Human conversation has a 200–500ms turn-taking window. Cross that and the user knows they&apos;re talking to a machine. A concrete production stack — Realtime API, parallel tool calls, predictive VAD — that hits sub-500ms on most turns, and where the latency budget actually goes.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Evals are your most important PM artifact</title>
      <link>https://refactor4ai.com/blog/evals-most-important-pm-artifact</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/evals-most-important-pm-artifact</guid>
      <pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate>
      <description>If you ship AI features and you don&apos;t own the eval suite, you don&apos;t really own the product. A practical guide to the eval craft AI PMs cannot delegate — what to write, how to structure it, and the failure modes that surface when you don&apos;t.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Embeddings for analysts (no ML required)</title>
      <link>https://refactor4ai.com/blog/embeddings-for-analysts</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/embeddings-for-analysts</guid>
      <pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate>
      <description>If you can write SQL you can use embeddings. A practical guide to using pgvector or DuckDB-VSS to cluster, search, and dedupe text data — without learning PyTorch or hiring an ML engineer. Real queries, real costs, things that surprised us.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Document AI: where OCR + LLM + structured output wins</title>
      <link>https://refactor4ai.com/blog/document-ai-ocr-llm-structured</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/document-ai-ocr-llm-structured</guid>
      <pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate>
      <description>Don&apos;t replace OCR with an LLM and don&apos;t pipe raw OCR into a chat model. The 2026 winning architecture is a hybrid — deterministic OCR for layout and high-volume extraction, multimodal LLMs for edge cases, structured-output schemas to bolt it together. Cost and accuracy numbers, with real pricing.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Cost Case: How a B2B support SaaS team cut their AI bill 68%</title>
      <link>https://refactor4ai.com/blog/cost-case-001-b2b-support-saas-cut-bill-68</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/cost-case-001-b2b-support-saas-cut-bill-68</guid>
      <pubDate>Fri, 15 May 2026 00:00:00 GMT</pubDate>
      <description>Cost Case Friday — B2B Customer Support SaaS, ~50 engineers, 620k AI-handled conversations/month. A team running every turn of a multi-turn support agent through a frontier model took their LLM bill from ~$187K/month to ~$59K/month in six weeks. Cost-aware cascade routing, semantic caching, prompt prefix collapse, output capping, self-hosted retrieval, and tighter top-K — with the measured contribution of each lever, the architecture before and after, and what they wish they&apos;d done first.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Token budgeting: a senior engineer&apos;s mental model</title>
      <link>https://refactor4ai.com/blog/token-budgeting-mental-model</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/token-budgeting-mental-model</guid>
      <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
      <description>A working mental model for budgeting LLM tokens — how to estimate per-request cost in your head, where the 10× overruns hide, and the four-number per-feature budget that prevents the end-of-month bill surprise. Concrete numbers for Claude 4.7, GPT-5.5, and Gemini 3.1 as of May 2026.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Postmortem #1: how an embedding model swap silently broke our retrieval for nine days</title>
      <link>https://refactor4ai.com/blog/postmortem-001-embedding-swap-broke-rag</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/postmortem-001-embedding-swap-broke-rag</guid>
      <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
      <description>Production Postmortem — EdTech, ~150 engineers. A routine embedding model upgrade landed in production with the index never re-embedded. Hybrid search rankings collapsed, the assistant began confidently citing the wrong source documents, retry-storms ran up a ~$48K LLM bill, and nobody noticed for nine days because no offline eval was wired to the retrieval layer. Timeline, root cause, fixes, and five lessons for anyone running RAG in production.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The Living Spec: markdown PRDs that ship</title>
      <link>https://refactor4ai.com/blog/living-spec-markdown-prds</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/living-spec-markdown-prds</guid>
      <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
      <description>Ditch the 20-page Google Doc PRD. The Living Spec pattern — a single markdown file checked into the repo next to the code, edited continuously by both PM and AI agents — is how AI-native teams keep documentation aligned with what&apos;s actually being built. A senior-PM playbook for the format, the rituals, and the trade-offs.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The &apos;lethal trifecta&apos; of agent security</title>
      <link>https://refactor4ai.com/blog/lethal-trifecta-agent-security</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/lethal-trifecta-agent-security</guid>
      <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
      <description>Simon Willison&apos;s lethal trifecta — private data, untrusted content, and an exfiltration vector — is the single most important security model in 2026. A senior engineer&apos;s guide to why prompt-injection defences fail, why structural mitigations are the only ones that work, and which leg of the trifecta is cheapest to cut in your architecture.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>FinOps for AI: cutting bills 40% in a quarter</title>
      <link>https://refactor4ai.com/blog/finops-for-ai-cutting-bills</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/finops-for-ai-cutting-bills</guid>
      <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
      <description>The eight-lever playbook we use to take AI bills down 40% in a quarter without degrading quality — prompt caching, batch APIs, routing, smaller models, capacity commitments, output capping, retrieval pruning, and the dashboard wiring that makes the savings stick. Concrete numbers on Claude, GPT-5, and Gemini for mid-2026.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Bedrock vs Azure OpenAI vs Vertex AI: a 2026 decision tree</title>
      <link>https://refactor4ai.com/blog/bedrock-vs-azure-openai-vs-vertex-ai-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/bedrock-vs-azure-openai-vs-vertex-ai-2026</guid>
      <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
      <description>An opinionated, decision-tree comparison of AWS Bedrock, Microsoft Azure AI Foundry, and Google Vertex AI as of May 2026 — which models live where, real list prices, throughput economics, the data-egress gotchas nobody puts in the slide deck, and the four questions that should actually pick your cloud.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Why your agents fail in production (and the 5 fixes that actually work)</title>
      <link>https://refactor4ai.com/blog/why-agents-fail-in-production</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/why-agents-fail-in-production</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>Most enterprise AI agents in 2026 never reach production — Deloitte&apos;s 2026 Tech Trends report puts production deployment at 11% even with 38% of organisations actively piloting. Five concrete failure modes show up in every post-mortem we run. Tool error handling, context drift, dumb RAG, brittle connectors, and no evals. Here&apos;s what each looks like and how to fix it.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Staffing AI features without a dedicated AI team</title>
      <link>https://refactor4ai.com/blog/staffing-ai-no-dedicated-team</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/staffing-ai-no-dedicated-team</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>You have been asked to ship AI features and you do not have ML engineers. A pragmatic 2026 staffing playbook for Staff+ engineers and EMs — what roles you actually need (and which ones you do not), how to upskill existing engineers in 90 days, and when to reach for fractional help instead of hiring.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Reranking: the single highest-ROI improvement you can make to a RAG pipeline</title>
      <link>https://refactor4ai.com/blog/reranking-highest-roi-rag</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/reranking-highest-roi-rag</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>If your RAG system retrieves 50 chunks and stuffs the top-5 into the prompt, you&apos;re shipping mediocre answers. A cross-encoder reranker on top of hybrid retrieval lifts retrieval accuracy 15–40% for 80–150ms of added latency. Here&apos;s the case for it, the production options in 2026, and the rerank-or-die patterns.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>RAG vs fine-tuning in 2026: a decision framework that actually ships</title>
      <link>https://refactor4ai.com/blog/rag-vs-fine-tuning-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/rag-vs-fine-tuning-2026</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>The RAG-vs-fine-tuning debate is mostly noise in 2026. The real question is where to place knowledge, where to encode behaviour, and how to evaluate both. Here&apos;s the decision framework, the order of operations, and the LoRA-on-top-of-retrieval pattern most teams should default to.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Writing PRDs for AI products: from deterministic to probabilistic</title>
      <link>https://refactor4ai.com/blog/prds-for-ai-products-probabilistic</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/prds-for-ai-products-probabilistic</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>Classic PRD templates assume a system that returns the same answer twice. AI products do not. A working PM template for 2026 — acceptance criteria as distributions, eval datasets as spec, failure modes as first-class requirements, and the living document model that survives a model upgrade.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>LLM observability: the 2026 stack you actually need</title>
      <link>https://refactor4ai.com/blog/llm-observability-stack-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/llm-observability-stack-2026</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>Stop flying blind on production AI. A practical 2026 reference stack for tracing, evals, cost analytics, and prompt management — built around Langfuse, Helicone, and the OpenTelemetry GenAI conventions — with the integration tradeoffs we actually run into in production.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Hybrid search 101: BM25 + vector + RRF, and why pure semantic search is leaving recall on the floor</title>
      <link>https://refactor4ai.com/blog/hybrid-search-bm25-vector-rrf</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/hybrid-search-bm25-vector-rrf</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>Vector-only retrieval is the wrong default in 2026. Pure semantic search tops out around 65–78% recall@10 on real-world RAG corpora. Hybrid retrieval — BM25 plus vector plus Reciprocal Rank Fusion — pushes that to 91%, and every major vector DB ships it natively. Here&apos;s exactly how it works and how to wire it up.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Computer-use agents 101: Claude, Operator, ACE</title>
      <link>https://refactor4ai.com/blog/computer-use-agents-101</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/computer-use-agents-101</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>A senior-engineer primer on computer-use agents in May 2026 — Claude Computer Use vs OpenAI Operator (CUA) vs General Agents&apos;s ACE. How each one drives a GUI, where they fail, the realistic accuracy you should plan for, and which is the right pick for which job.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Building an AI gateway: rate limits, retries, fallback</title>
      <link>https://refactor4ai.com/blog/ai-gateway-rate-limits-fallback</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-gateway-rate-limits-fallback</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>Direct vendor calls are a single point of failure. A senior-engineer blueprint for the AI gateway pattern in 2026 — virtual keys, per-tenant rate limits, graceful retry with backoff, cross-vendor fallback, prompt caching, and the buy-vs-build choice between LiteLLM, Portkey, Kong AI, and rolling your own.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Agent loops in 2026: ReAct vs Plan-and-Execute vs Reflection</title>
      <link>https://refactor4ai.com/blog/agent-loops-react-vs-plan-execute</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/agent-loops-react-vs-plan-execute</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate>
      <description>Most production agents in 2026 aren&apos;t pure ReAct or pure Plan-and-Execute — they&apos;re a hybrid that picks each pattern for the part of the task it&apos;s good at. Here&apos;s what the three loops actually do, the cost-latency-quality tradeoff, and the architecture pattern that wins for long-running workflows.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>What is MCP? A practical guide to the Model Context Protocol in 2026</title>
      <link>https://refactor4ai.com/blog/what-is-mcp-model-context-protocol</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/what-is-mcp-model-context-protocol</guid>
      <pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate>
      <description>MCP is the standard way LLMs talk to tools, databases and APIs in 2026. This is the plain-English explainer — what MCP is, why it won, and how to build your first MCP server in TypeScript or Python.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>Prompt caching: how to cut your LLM bill by 90% in 2026</title>
      <link>https://refactor4ai.com/blog/prompt-caching-90-percent-savings</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/prompt-caching-90-percent-savings</guid>
      <pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate>
      <description>Prompt caching is the single biggest lever for reducing LLM costs in production. Here&apos;s how it works on Anthropic, OpenAI and Google, exactly what to cache, and the patterns that turn a $30k/day feature into a $3k/day one.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The 2026 AI System Design Interview: a complete preparation guide</title>
      <link>https://refactor4ai.com/blog/ai-system-design-interview-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-system-design-interview-2026</guid>
      <pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate>
      <description>AI system design is the new system design interview. Here&apos;s the format, the categories of questions FAANG and AI labs are actually asking, the framework to structure an answer, and the trade-offs you&apos;ll be expected to articulate.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>AI fluency for product managers: what to actually learn in 2026</title>
      <link>https://refactor4ai.com/blog/ai-fluency-for-product-managers</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-fluency-for-product-managers</guid>
      <pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate>
      <description>AI fluency for PMs isn&apos;t &apos;know what an LLM is.&apos; It&apos;s writing specs engineers can build, owning the eval plan, classifying risk under the EU AI Act, and pricing AI features. Here&apos;s the practical curriculum.</description>
      <author>Refactor4AI team</author>
    </item>
    <item>
      <title>The 2026 AI Capability Map: Claude Opus 4.7 vs GPT-5.5 vs Gemini 3.1 Pro</title>
      <link>https://refactor4ai.com/blog/ai-capability-map-2026</link>
      <guid isPermaLink="true">https://refactor4ai.com/blog/ai-capability-map-2026</guid>
      <pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate>
      <description>A practical, role-agnostic capability map for the May 2026 flagship models — costs per million tokens, effective context windows, reasoning premiums, multimodal coverage, cloud availability, and which to actually pick for which job.</description>
      <author>Refactor4AI team</author>
    </item>
  </channel>
</rss>
