<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Engineering — Felesh</title><description>Engineering</description><link>https://blog.felesh.ai/</link><item><title>Framework or platform? Choosing the right abstraction for your agent system</title><link>https://blog.felesh.ai/doors/framework-or-platform/</link><guid isPermaLink="true">https://blog.felesh.ai/doors/framework-or-platform/</guid><description>Past the prototype, you are assembling your third agent by hand — and the coordination tax is real.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Off-the-shelf, or adapt your own?</title><link>https://blog.felesh.ai/doors/off-the-shelf-or-adapt-your-own/</link><guid isPermaLink="true">https://blog.felesh.ai/doors/off-the-shelf-or-adapt-your-own/</guid><description>The demo worked. Your production domain did not. Before you commit to training your own model, understand the adaptation ladder and why most teams stop far lower than they think they need to.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate></item><item><title>When one agent isn&apos;t enough</title><link>https://blog.felesh.ai/rail-builder/many-agents-to-cwd/</link><guid isPermaLink="true">https://blog.felesh.ai/rail-builder/many-agents-to-cwd/</guid><description>When a single agent buckles under the weight of too many responsibilities, you have to split the work — without losing the shared context.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Run it for real</title><link>https://blog.felesh.ai/rail-builder/run-it-for-real/</link><guid isPermaLink="true">https://blog.felesh.ai/rail-builder/run-it-for-real/</guid><description>Moving from local prototype to production serving — how to select, adapt, and run models once cost, latency, and reliability actually matter.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The agent that acts (and when not to)</title><link>https://blog.felesh.ai/rail-builder/the-agent-that-acts/</link><guid isPermaLink="true">https://blog.felesh.ai/rail-builder/the-agent-that-acts/</guid><description>An LLM answers questions; an agent takes action. Knowing when to hand over control — and when to hold it back — decides whether the system stays stable.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The augmented LLM</title><link>https://blog.felesh.ai/rail-builder/the-augmented-llm/</link><guid isPermaLink="true">https://blog.felesh.ai/rail-builder/the-augmented-llm/</guid><description>A raw prompt in a playground creates an illusion. In production the working unit is an augmented model — wrapped with memory, tools, and context.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The LLM is a component</title><link>https://blog.felesh.ai/rail-builder/the-llm-is-a-component/</link><guid isPermaLink="true">https://blog.felesh.ai/rail-builder/the-llm-is-a-component/</guid><description>A language model alone is a raw text engine; every capability — memory, action, reliability — has to be engineered around it.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The Coordinator, the Worker, and the Delegator</title><link>https://blog.felesh.ai/agent-architecture/cwd-anchor/</link><guid isPermaLink="true">https://blog.felesh.ai/agent-architecture/cwd-anchor/</guid><description>One agent can only carry so much. When you split the work across many, you need a shape for who decides, who does, and who keeps the whole thing balanced. That shape has a name.</description><pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Three lessons the human mind offers agent architecture</title><link>https://blog.felesh.ai/agent-architecture/brain-lessons-for-agent-architecture/</link><guid isPermaLink="true">https://blog.felesh.ai/agent-architecture/brain-lessons-for-agent-architecture/</guid><description>Several hard problems in agent design rhyme with how the human mind works: focus over clutter, knowing when a job is done, and separating the layers of memory. The resemblances are a good guide for design.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Cognitive Least Privilege: your agent should know only what it needs</title><link>https://blog.felesh.ai/agent-architecture/cognitive-least-privilege/</link><guid isPermaLink="true">https://blog.felesh.ai/agent-architecture/cognitive-least-privilege/</guid><description>Any information that doesn&apos;t serve the agent&apos;s job both lowers accuracy and widens the attack surface. Borrow least privilege from security and extend it to what an agent knows.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Evaluate models on your own set, not the public leaderboard</title><link>https://blog.felesh.ai/model-selection/evaluate-on-your-own-set/</link><guid isPermaLink="true">https://blog.felesh.ai/model-selection/evaluate-on-your-own-set/</guid><description>Public leaderboards tell you less about your work than you think. The reliable method: build a small eval set that represents your real task, and score the candidates on that.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Event-driven by design: agent teams that don&apos;t lose messages</title><link>https://blog.felesh.ai/agent-architecture/event-driven-agent-teams/</link><guid isPermaLink="true">https://blog.felesh.ai/agent-architecture/event-driven-agent-teams/</guid><description>When several agents work together, the biggest risk is lost messages and a collapsing chain. An event-driven architecture removes that risk with a few simple rules.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Fine-tune, RAG, or prompt: which one, and what each costs</title><link>https://blog.felesh.ai/fine-tuning/fine-tune-rag-or-prompt/</link><guid isPermaLink="true">https://blog.felesh.ai/fine-tuning/fine-tune-rag-or-prompt/</guid><description>There are three ways to adapt a model to your need, and the wrong choice can get expensive. The difference is in what problem each one actually solves.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Common fine-tuning pitfalls and how to debug them</title><link>https://blog.felesh.ai/fine-tuning/fine-tuning-pitfalls/</link><guid isPermaLink="true">https://blog.felesh.ai/fine-tuning/fine-tuning-pitfalls/</guid><description>Most failed fine-tunes trace back to a few recurring patterns. If you know the signs, debugging becomes a simple checklist instead of guesswork.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>From prompt engineering to context engineering</title><link>https://blog.felesh.ai/prompting/from-prompt-to-context-engineering/</link><guid isPermaLink="true">https://blog.felesh.ai/prompting/from-prompt-to-context-engineering/</guid><description>There was a time when the art of working with a model came down to writing a good prompt. But the centre of gravity is shifting: from crafting one instruction to designing the whole context the model works in.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>The LoRA family: QLoRA, DoRA, and LoRA+ — which, and when?</title><link>https://blog.felesh.ai/fine-tuning/lora-family-variants/</link><guid isPermaLink="true">https://blog.felesh.ai/fine-tuning/lora-family-variants/</guid><description>Since LoRA was introduced, several improved variants have appeared, each targeting one particular problem. Knowing them helps you pick the right one for each job.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>LoRA hyperparameters demystified: rank, alpha, and what to set</title><link>https://blog.felesh.ai/fine-tuning/lora-hyperparameters-demystified/</link><guid isPermaLink="true">https://blog.felesh.ai/fine-tuning/lora-hyperparameters-demystified/</guid><description>Fine-tuning with LoRA has a handful of key numbers, and once you know what they mean, choosing them is simple. This guide clears up rank, alpha, learning rate, and the rest.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>MLP is the model&apos;s memory: where knowledge lives</title><link>https://blog.felesh.ai/llm-infra/mlp-is-the-models-memory/</link><guid isPermaLink="true">https://blog.felesh.ai/llm-infra/mlp-is-the-models-memory/</guid><description>In a language model, attention layers route information, but the actual knowledge is stored elsewhere — in the MLP layers that make up the bulk of the model.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>PagedAttention and continuous batching: how one server answers more users</title><link>https://blog.felesh.ai/llm-infra/paged-attention-continuous-batching/</link><guid isPermaLink="true">https://blog.felesh.ai/llm-infra/paged-attention-continuous-batching/</guid><description>Two infrastructure tricks multiply the capacity of a language-model server: continuous batching and smart KV-cache management. Both come from one simple idea — don&apos;t waste resources.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>A practical checklist for picking an LLM for your feature</title><link>https://blog.felesh.ai/model-selection/pick-an-llm-checklist/</link><guid isPermaLink="true">https://blog.felesh.ai/model-selection/pick-an-llm-checklist/</guid><description>Choosing a model is less about leaderboards than about knowing your own need precisely. Six simple steps you can follow today.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>How LLM inference actually works: prefill vs decode</title><link>https://blog.felesh.ai/llm-infra/prefill-vs-decode/</link><guid isPermaLink="true">https://blog.felesh.ai/llm-infra/prefill-vs-decode/</guid><description>Text generation has two phases with very different behaviour: one compute-bound, one memory-bound. Understanding the difference explains why the KV cache exists and why decode is slow.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>When no single model is enough: the primary and verifier pattern</title><link>https://blog.felesh.ai/model-selection/primary-and-verifier/</link><guid isPermaLink="true">https://blog.felesh.ai/model-selection/primary-and-verifier/</guid><description>Sometimes a task has two critical demands that no single model satisfies at once. The answer isn&apos;t to accept a weak model; it&apos;s to combine two.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Common failure modes in LLM systems — and how to catch them</title><link>https://blog.felesh.ai/model-selection/production-failure-modes/</link><guid isPermaLink="true">https://blog.felesh.ai/model-selection/production-failure-modes/</guid><description>A language model fails in specific ways, not random ones. If you know these modes, you can catch them before your users do.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Stop ranking LLMs, start profiling them</title><link>https://blog.felesh.ai/model-selection/profile-dont-rank-llms/</link><guid isPermaLink="true">https://blog.felesh.ai/model-selection/profile-dont-rank-llms/</guid><description>A single number on a leaderboard won&apos;t tell you which model fits your job. A multi-dimensional profile of capabilities will.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Defending against prompt injection and jailbreaks — and reducing hallucination</title><link>https://blog.felesh.ai/agent-architecture/prompt-injection-and-defense/</link><guid isPermaLink="true">https://blog.felesh.ai/agent-architecture/prompt-injection-and-defense/</guid><description>When user input can change an agent&apos;s behaviour, security becomes a design problem. A few clear principles neutralise most of these attacks.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Fine-tune your first model on free Colab: QLoRA in about 40 lines</title><link>https://blog.felesh.ai/fine-tuning/qlora-on-free-colab/</link><guid isPermaLink="true">https://blog.felesh.ai/fine-tuning/qlora-on-free-colab/</guid><description>Fine-tuning a model doesn&apos;t have to need an expensive cluster. With QLoRA you can tune a small model on a free GPU in just a few dozen lines of code.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>A 70B model on one GPU: a practical guide to quantization</title><link>https://blog.felesh.ai/llm-infra/run-a-70b-on-one-gpu/</link><guid isPermaLink="true">https://blog.felesh.ai/llm-infra/run-a-70b-on-one-gpu/</guid><description>A seventy-billion-parameter model needs about 140 GB of memory at full precision. With quantization you can compress that same model until it fits on a single GPU — and keep quality almost untouched.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Save first, then publish: a simple rule for not losing work</title><link>https://blog.felesh.ai/agent-architecture/save-before-publish/</link><guid isPermaLink="true">https://blog.felesh.ai/agent-architecture/save-before-publish/</guid><description>One of the most common hidden bugs in event-driven systems is publishing the news before the fact is recorded. The right order — save first, then publish — removes that bug at the root.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>The rule layer: deterministic guardrails around a probabilistic model</title><link>https://blog.felesh.ai/model-selection/the-rule-layer/</link><guid isPermaLink="true">https://blog.felesh.ai/model-selection/the-rule-layer/</guid><description>A language model is probabilistic and sometimes errs. The way to make it reliable isn&apos;t to perfect the model; it&apos;s to build a deterministic layer that catches what slips through.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Tracing a request through a multi-agent system</title><link>https://blog.felesh.ai/agent-architecture/trace-a-request-through-agents/</link><guid isPermaLink="true">https://blog.felesh.ai/agent-architecture/trace-a-request-through-agents/</guid><description>The best way to understand a multi-agent architecture is to follow a real request from start to finish. Let&apos;s trace a vague message, step by step, into a structured action.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>What quantization actually does: precision loss and vector-space collapse</title><link>https://blog.felesh.ai/llm-infra/what-quantization-actually-does/</link><guid isPermaLink="true">https://blog.felesh.ai/llm-infra/what-quantization-actually-does/</guid><description>Quantization means holding a model&apos;s weights with fewer bits. But what exactly does this loss of precision do to the model, and why are models so robust to it?</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Where LLM serving costs actually go</title><link>https://blog.felesh.ai/llm-infra/where-llm-serving-costs-go/</link><guid isPermaLink="true">https://blog.felesh.ai/llm-infra/where-llm-serving-costs-go/</guid><description>If you open up the bill for serving a model, most of the cost is concentrated in one place. Understanding that concentration also clarifies the eternal &apos;build or buy&apos; question.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Why LoRA works: the intrinsic-dimensionality story</title><link>https://blog.felesh.ai/fine-tuning/why-lora-works/</link><guid isPermaLink="true">https://blog.felesh.ai/fine-tuning/why-lora-works/</guid><description>If a large model has billions of parameters, how can you tune it by training only a few small matrices? The answer is a subtle idea: the change you need has a small intrinsic dimension.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>From worker to specialist: an agent that owns a domain</title><link>https://blog.felesh.ai/agent-architecture/worker-to-specialist/</link><guid isPermaLink="true">https://blog.felesh.ai/agent-architecture/worker-to-specialist/</guid><description>The difference between an executing worker and a specialist is that the first does one job and steps aside, while the second owns a domain and carries its state over time.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Working with LLM APIs: first calls, tokens, and structured output</title><link>https://blog.felesh.ai/prompting/working-with-llm-apis/</link><guid isPermaLink="true">https://blog.felesh.ai/prompting/working-with-llm-apis/</guid><description>Your first call to a language-model API is simpler than it looks. Once you know a few basics — roles, tokens, temperature, and structured output — the rest falls into place.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Zero-shot, few-shot, or chain-of-thought: picking the right technique</title><link>https://blog.felesh.ai/prompting/zero-shot-few-shot-cot/</link><guid isPermaLink="true">https://blog.felesh.ai/prompting/zero-shot-few-shot-cot/</guid><description>There are three basic prompting techniques, and each has its place. Knowing when to reach for which matters more than the techniques themselves.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate></item></channel></rss>