<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://sig-intent.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://sig-intent.com/" rel="alternate" type="text/html" hreflang="en" /><updated>2026-08-19T10:58:11+00:00</updated><id>https://sig-intent.com/feed.xml</id><title type="html">SIGINTENT | Writing</title><subtitle>Independent technology advisory by Antonio Elena. Fractional CTO, cloud, AI strategy and software architecture for founders, CTOs and executive teams facing the decisions that compound.</subtitle><author><name>Antonio Elena</name></author><entry><title type="html">The evolution of the software architect: from builder to system designer</title><link href="https://sig-intent.com/writing/evolution-of-the-software-architect/" rel="alternate" type="text/html" title="The evolution of the software architect: from builder to system designer" /><published>2026-08-19T00:00:00+00:00</published><updated>2026-08-19T00:00:00+00:00</updated><id>https://sig-intent.com/writing/evolution-of-the-software-architect</id><content type="html" xml:base="https://sig-intent.com/writing/evolution-of-the-software-architect/"><![CDATA[<p>The first time I was called a software architect, the job was to draw box diagrams.</p>

<p>This was the mid-2000s. The role was well-defined in the consultant playbook of the era: you were the person who produced the UML, picked the design patterns, chose between WCF and REST, and wrote the standards document that the development team was supposed to follow. Your deliverable was a folder of Visio files no one really paid any attention to, and some documents with the word “framework” in the titles.</p>

<p>I did this job for several years. I did not particularly enjoy all of it. The parts I found interesting — the conversations with the business, the architectural nuances and choices, the trade-off and priority reasoning, the political negotiation about what should be built and why — were not officially part of the role. They were side effects but the essence of it too. The deliverables did not really matter that much even when the role, as written, was about producing the diagrams, because someone somewhere is footing an expensive bill. Everything else was tolerated as long as you kept producing the diagrams. But it was probably the other way around.</p>

<p>Twenty years later, the role has changed in ways that the title has not caught up with. The diagrams still exist somewhere, but they are no longer the job. The new job is something else, and it is worth being explicit about what.</p>

<h2 id="what-the-old-role-was-actually-optimizing-for">What the old role was actually optimizing for</h2>

<p>The builder-architect of the 2000s and early 2010s was optimizing for something real: <strong>reducing variance in large codebases</strong> (aka chaos and drift and mud). In a world where dozens or even hundreds of developers of varying skill levels were working on the same system, and where the cost of a bad abstraction was measured - and not that often - in years of maintenance, the architect’s job was to impose structure from above, even when misguided or too constrained. You picked the patterns, often in advance and with little information, applying recipes and muscle memory from past engagements, and wins and defeats. You wrote the standards. You enforced the conventions. The goal was not to make the codebase brilliant; it was to make the codebase as <em>consistent</em> as possible, because some degree of consistency was the only way to keep a large team from producing a mess. Even when it was largely an uphill battle in the face of constant pressure to deliver more faster and cut whatever corners in the process.</p>

<p>This worked, mostly, in varying degrees, for the problems of its time. It produced legible enterprise systems and did not prevent showcase textbook failures. It gave junior developers a scaffolding to work inside (some constraint and structure is actually liberating). It kept the variance down. It was necessary.</p>

<p>It was also deeply, structurally conservative. The architect’s role was to reduce risk, and reducing risk meant saying no to most new ideas. The failure mode of the builder-architect was not a bad system — it was a system that could not change, because every change had to pass through the architect’s filter, and the filter got narrower with every year of service. By the mid-2010s, the term “architect” had become mildly derogatory in many engineering circles, as well as on the business side of the divide. It meant the person who was slowing you down. Some organizations kept these weird folk around, in their particular ivory tower, a relic from the past they did not dare tear down for fear some obscure function would cease working and something ignote would break. Or just out of inertia or reputation.</p>

<p>That was not the architect’s fault. It was the role’s fault. The role was designed to produce conservatism, and conservatism is what it produced.</p>

<h2 id="what-the-new-role-is-optimizing-for">What the new role is optimizing for</h2>

<p>The system-designer architect is optimizing for something different: <strong>making good decisions at the seams</strong>. Not inside individual systems, which modern tooling and modern teams can usually handle on their own, but at the seams where systems, teams, products, and strategies meet and have to agree on something.</p>

<p>The work has moved outward, in two directions.</p>

<p>Outward into the business. The modern architect spends meaningful time with product leaders, business heads, finance people, and occasionally the board. Not because they have become executives — most of them are still hands-on technically — but because the decisions that matter for the architecture are increasingly decisions about what the business is trying to do, and those decisions are not made in the engineering org.</p>

<p>Outward into the seams. Between cloud and on-premise. Between platforms and products. Between ML and distributed systems. Between vendors and in-house. Between the architecture you inherited and the architecture the new demands that AI features want. These are the places no single specialist can own, which is exactly why the interesting problems live there.</p>

<p>The old architect produced diagrams of systems, largely self-contained with little or no outside calls and dependencies. The new architect’s deliverable is not a static Visio file anymore. It is a crisp, defensible, live set of documented decisions with their assumptions written down, reversibility classified, optionality built-in and its implications explained to the people who need to understand them in a language that they understand (this has not really changed since then and communication is still key).</p>

<p>Gregor Hohpe named this movement better than I will. His architect elevator rides between the penthouse and the engine room because the people on each floor cannot hear each other, and somebody has to carry the message without it turning into the telephone game. That destination has been described. What is missing is what the ride costs, and what happens when the building is not a skyscraper.</p>

<h2 id="three-things-the-elevator-does-not-tell-you">Three things the elevator does not tell you</h2>

<p><strong>Standing decays.</strong> The elevator assumes you are allowed on it. Credibility on the engine-room floor is not a title, it is recent evidence, and it has a half-life. An architect who has not shipped anything in four years is visiting, not riding, and everyone on that floor can tell inside ten minutes. The decay is invisible from the penthouse, where the same architect still sounds authoritative, which is how you end up with someone who has lost one of their two floors and does not know it.</p>

<p><strong>Most buildings have three floors, not thirty.</strong> Hohpe’s frame is the large enterprise and he says so: a skyscraper so tall that one elevator might not span it. Half the companies I work with have forty people. The CTO is in the engine room and in front of the board the same morning. There is no distance to cover, so the skill is not travel at all. It is holding both altitudes at once, in one conversation, without changing register when the founder walks past. The enterprise literature barely names that, because the enterprise never has to do it.</p>

<p><strong>Increasingly, the architect is not an employee.</strong> The elevator is internal: it presumes a badge, a reporting line, and a building you belong to. The advisory architect has none of that. You do not ride, you are invited to a floor, once, and you get one meeting to establish that you belong on it. Nothing accrues between engagements. That entry cost is the largest hidden expense in the model, for you and for the client paying you to spend two weeks earning the right to be listened to.</p>

<h2 id="the-skills-that-matter-now">The skills that matter now</h2>

<p>Four things, and the fourth is what makes the other three count for anything.</p>

<p><strong>Framing.</strong> The ability to take a vague, tangled, multi-stakeholder question and reformulate it as a specific decision with named options and named tradeoffs. This is the single hardest skill in the new role and the one that most differentiates senior architects from junior ones. A staff engineer can answer a well-framed question. A principal architect reformulates the question so that it can be answered at all.</p>

<p><strong>Translation.</strong> The ability to move fluently between the vocabulary of engineers, product managers, executives, and business stakeholders — and, critically, to do it in real time, in the same conversation, without code-switching awkwardly. The architect’s job now includes explaining to a CEO why a particular architectural decision can affect a critical process, some competitive advantage or the bottom line down the line in eighteen months, and doing it in language the CEO can act on. And to do this without sounding ominous or pessimistic, even when the point might be somber and still needs to be driven home. This is human work, translation work, not technical work, and it is where most senior architects are undertrained, and scars can only be gained in the battle field.</p>

<p><strong>Judgment about reversibility.</strong> The architect’s most valuable output is knowing how permanent each decision is, and making sure the team treats it with the right weights in mind. This is not glamorous but rugged. It is also most of the job. Context, nuance, long-term vision, Wardley Maps and other tools, you have to hold in mind much more than others do, can or are even willing to, in their narrower tactical domains.</p>

<p><strong>Having to live with it.</strong> Everything above describes producing decisions, and the obvious objection is that anyone can write a memo. Hohpe calls the failure mode authority without responsibility, and it breaks only when the architect has to live with the consequences, or at least stay close enough to watch them land. An architect who frames a decision, documents it, hands it over and leaves before the bill arrives has made a recommendation, not a decision, and a recommendation is a cheaper thing. The distinction is not moral but epistemic: you do not find out whether your judgment was any good unless you are still there when the system tells you.</p>

<p>Notice what is not on this list. Pattern catalogs. UML. Framework design guidelines. These are not irrelevant — they are still part of the toolkit — but they are not where senior architects spend most of their time anymore. They are table stakes, not differentiators.</p>

<h2 id="the-personal-part">The personal part</h2>

<p>I will be honest: the shift in the role has been good for me, because the parts of the job I always found interesting — the framing, the negotiations, the cross-domain mediations — are now the parts the role is measured on. The parts I found less creative with time, once the novelty of the junior architect wears off — the diagrams, the standards documents, the governance reviews — have become less central. I do not miss them.</p>

<p>But major shifts are rarely kind to everyone. Architects who built their careers around being the person who boasted the technical prowess, who knew the most patterns, or the person who had the purview to write and enforce the standards, or the person who produced the best diagrams, have found themselves in a role they did not sign up for. The role asks for skills they were never hired for or developed. Some of them have adapted. Some of them perhaps have not that well.</p>

<p>The ones who adapted did so by doing a specific thing: they stopped thinking of themselves as builders of systems and started thinking of themselves as designers of decisions, expert guides in a complicated landmined terrain no one has the entire map to. Once that reframe lands, the rest is learnable. Before it lands, the new role feels like an imposition to redefine our identity. After it lands, the old role feels like an old cage.</p>

<h2 id="the-last-ten-things-you-made">The last ten things you made</h2>

<p>If you are a software architect wondering whether your role is heading somewhere you want to go, try this: look at the last ten things you produced that were genuinely useful to your organization. How many of them were diagrams? How many of them were documented decisions with framing, tradeoffs, and implications, ADRs perhaps? How many were actually not that tangible even?</p>

<p>If most of them were diagrams, you are still doing the old job. That is fine, and it is sometimes still the right job — but it is not where the senior work is going. If most of them were documented decisions, you are doing the new job, whether your title has caught up or not. The title will eventually catch up. The work is what you defend.</p>

<p>And then ask the harder one, which is the fourth skill wearing everyday clothes: how many of those ten did you stay to see the consequences of?</p>

<p>The architect used to build systems. Now the architect designs the decisions that determine what the systems become. Both are real. Only one is what the role is becoming.</p>]]></content><author><name>Antonio Elena</name></author><category term="architecture" /><category term="careers" /><category term="engineering-org" /><category term="decision-craft" /><summary type="html"><![CDATA[The architect's job moved from drawing diagrams to framing decisions. What the elevator model leaves out is what the ride costs to keep.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sig-intent.com/assets/img/og-default.png" /><media:content medium="image" url="https://sig-intent.com/assets/img/og-default.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">M-shaped engineering: cheap breadth is why you need more than one depth</title><link href="https://sig-intent.com/writing/m-shaped-engineering/" rel="alternate" type="text/html" title="M-shaped engineering: cheap breadth is why you need more than one depth" /><published>2026-08-17T00:00:00+00:00</published><updated>2026-08-17T00:00:00+00:00</updated><id>https://sig-intent.com/writing/m-shaped-engineering</id><content type="html" xml:base="https://sig-intent.com/writing/m-shaped-engineering/"><![CDATA[<p>For twenty years the industry told engineers to specialize. Choose your niche and go deep, or so the story went. The advice was right. The conditions that made it right are gone.</p>

<p>T-shaped — broad awareness across many areas, real depth in one — matched the economics of technical work in an era when context was expensive. Getting to genuine competence in a new domain took years of reading, building, and being wrong in public: a considerable investment, and an emotional one, since a specialty that took a decade ends up commingled with self-identity. When acquisition costs that much, specializing is rational. It is also a bet, and a concentrated one. You go deep once, and the market pays you for having run the long distance.</p>

<p>Those economics broke somewhere in the last three years. But they did not break the way the “AI makes everyone a generalist” argument claims, and getting the distinction right is the difference between a useful career strategy and an expensive mistake.</p>

<p>What got cheap is <strong>access</strong> to a domain — or more precisely, good-enough-for-most-cases initial access. Judgment inside it did not get cheap at all.</p>

<h2 id="the-objection-which-is-correct">The objection, which is correct</h2>

<p>The strongest argument against everything I am about to say goes like this:</p>

<blockquote>
  <p><em>A generalist with a language model is a menace. The tool produces fluent, confident, structurally wrong answers in every domain simultaneously, and it produces them in exactly the register that passes for expertise. The person least equipped to catch this is the one with broad, shallow exposure — because shallow exposure gives you the vocabulary to follow the answer, threadbare knowledge to sanity-check it, and none of the scar tissue to doubt it. Depth is the only reliable antidote. Therefore depth is now</em> more <em>valuable, not less, and organizations should be hiring harder specialists rather than softer generalists.</em></p>
</blockquote>

<p>Every step of that is right except the last one.</p>

<p>It is right about the mechanism. Fluent-and-wrong is a characteristic failure mode of the current tooling, and shallow breadth is genuinely the worst possible defence against it. I have watched competent people ship bad decisions in the last eighteen months specifically because an answer arrived in a domain adjacent to theirs, sounded correct, and they had no basis on which to distrust it.</p>

<p>Where it goes wrong is the prescription — because the antidote it proposes is unbuildable. You cannot have depth in the domain where you are being fooled. That is what “being fooled” means. Nobody has depth in eight domains, and the failures do not politely occur inside the one you own.</p>

<h2 id="what-actually-transfers">What actually transfers</h2>

<p>The useful asset is not depth in the domain under discussion. It is <strong>having been deep at all, more than once.</strong></p>

<p>Here is the mechanism, and it is the reason the shape is M and not T.</p>

<p>Your first depth teaches you a domain. Your second depth teaches you something different and more portable: what the <em>transition</em> feels like — the specific sensation of moving from confident-and-wrong to actually-correct, of discovering that the clean mental model you held for two years was load-bearing in the wrong place. You only learn that by having held a wrong model long enough to be embarrassed by it, and then doing it again somewhere else, so that you recognise the pattern as a pattern rather than as one bad week.</p>

<p>That sensation is domain-independent. It is what fires when a fluent answer arrives in a field you do not own and something about its smoothness is wrong. The person with one depth has a single calibrated instrument and trusts the machine everywhere outside it. The person with two or three has learned what the inside of real competence feels like, which is precisely what lets them detect its absence in territory they have never worked.</p>

<p>So the case for M-shaped is not that breadth beats depth. Breadth is now free, and free things do not win arguments. The case is that the second and third depths are what make free breadth <strong>safe to use</strong>.</p>

<h3 id="a-note-on-the-letters">A note on the letters</h3>

<p>There is a taxonomy in circulation, mostly in the HR literature, that counts the verticals: T for one specialization, M for two or more, and Comb for many, each resting on a broad base of general skills. It is a useful vocabulary and I will use M throughout, because M is where the argument lives, but consider the fundamentally equivalent.</p>

<p>But the counting is the least interesting part of it, and I think it is actively misleading. The transformation is not linear in the number of depths. It happens <strong>once, at the second one</strong> — because the second depth is the first evidence you have that your first depth was a way of thinking rather than the way things are. A third and fourth add reach and range, and they compound in ways worth having, but they do not repeat that event. Nobody becomes twice as hard to fool by acquiring a fifth specialty.</p>

<p>Which means the comb profile is not the aspirational end of a ladder that starts at T. The ladder has exactly one rung that matters, and most people never step onto it.</p>

<h3 id="what-the-horizontal-bar-actually-is">What the horizontal bar actually is</h3>

<p>The taxonomy is right that the verticals rest on something, and wrong about what. The usual answer is soft skills: communication, empathy, teamwork. Those are good things to have and they are not the load-bearing element.</p>

<p>The base is <strong>translation</strong> — the ability to carry a constraint out of one domain’s vocabulary and into another’s without dropping it on the way. “This index will not survive the write volume” and “the month-end close will miss its window” are the same fact stated to two audiences, and someone has to be able to hold both forms of it at once and know they are the same. That capacity is what makes several depths compose into judgment instead of sitting alongside each other as unrelated party tricks.</p>

<p>Gregor Hohpe’s architect elevator makes the adjacent point about organizations: the value is in riding between the penthouse and the engine room, because the people on each floor cannot hear each other. What the elevator model understates is the entry condition. You cannot ride credibly to a floor whose language you have never actually spoken. Two or more depths are what buy you the ticket.</p>

<p>One thread I am deliberately leaving alone: depths that are <em>not</em> adjacent — that share little or no Venn overlap — produce the most interesting cross-pollination, and also the hardest translation problem. That deserves its own essay rather than a paragraph in this one.</p>

<h2 id="what-this-looks-like-at-scale">What this looks like at scale</h2>

<p>Running a global architecture function — 800+ projects, reporting into the Group CIO — gives you an unusual sample: not one organization’s decisions, but hundreds of them, taken by different teams under different pressures, with the outcomes visible two and three years later.</p>

<p>The decisions that went badly were almost never wrong <em>inside</em> a domain. Domain experts are good at their domains; the database people did not choose bad indexes and the network people did not misconfigure the peering. The failures clustered at the seams: where the data model met the integration pattern, where the cloud cost model met the internal chargeback rules, where the security posture met the delivery cadence the business had already promised to a customer.</p>

<p>Not that in-domain errors never compound together — they do. But they compound legibly, inside a single vocabulary, in front of someone whose job it is to notice.</p>

<p>Seams have no owner, almost by organizational design. Every owner was a specialist, and the seam belonged to none of them — so it was nobody’s job to notice that two locally correct decisions composed into one globally wrong system.</p>

<p>The structural conclusion is uncomfortable for anyone who has built an org chart: <strong>an organization staffed entirely with T-shaped people has a complete map with unowned borders.</strong> Every territory is covered. Every boundary is not.</p>

<h2 id="three-changes-and-what-each-one-costs">Three changes, and what each one costs</h2>

<p><strong>Stop treating depth as the default criterion for senior hires.</strong> Not “stop hiring specialists” — that would be a bad reading and an expensive mistake. You still need the person who can read a query planner at three in the morning, and if you staff entirely for breadth you will find out what that person was worth during your first genuine performance crisis, when nobody in the room has ever profiled anything. The change is narrower: for roles whose actual job is judgment under uncertainty, ask for two depths — or more — rather than for the deepest available one.</p>

<p><em>The cost:</em> your interview loop gets worse before it gets better. Depth is easy to test — you probe until the candidate runs out of answers. Second-depth is hard to test and easy to fake, and you will make more hiring mistakes for a year while you learn to tell transfer from tourism.</p>

<p><strong>Stop penalizing cross-domain transfer.</strong> Note the verb. The HR literature on these profiles reaches for development programmes — job rotation, cross-mentoring, personalised learning tracks — as though a second depth were something an organization can install in someone. It is not. The mechanism is holding a wrong model long enough to be embarrassed by it, and no rotation scheme manufactures that; six months on an adjacent team produces vocabulary, which is the failure mode, not the cure.</p>

<p>What an organization can actually do is stop punishing the people who do it anyway. Most review systems code cross-domain movement as “lack of focus” and quietly penalize the exact behaviour they claim to want.</p>

<p><em>The cost:</em> you will also stop penalizing some genuine dilettantes. The mitigation is a hard requirement — the transfer only counts if it produced a shipped outcome in the new domain. Exposure does not count. Curiosity does not count. Shipped counts.</p>

<p><strong>Stop making the single-specialty ladder the only senior ladder.</strong> The senior → staff → principal progression that rewards twenty years in one specialty still makes sense for a small number of real specialists. It stopped making sense as the default shape of a technical career.</p>

<p><em>The cost:</em> legibility. Single-domain ladders are easy to calibrate across a large organization; M-shaped ones are not, and levelling conversations get longer and more contested. Pay that cost knowingly rather than discovering it in the first calibration cycle.</p>

<h2 id="if-i-am-wrong">If I am wrong</h2>

<p>The bet has a real downside and I would rather name it than pretend otherwise.</p>

<p>If I am wrong, we spend a decade promoting confident generalists who have read about everything and been deep in nothing, the actual specialists conclude the ladder no longer rewards them and leave, and we find out what that cost the first time something breaks in a way no tool has seen before. That is not a small risk. Organizations are much better at destroying specialist career paths than at rebuilding them.</p>

<p>The hedge is not to pick one shape. It is to be deliberate about which roles want which — M-shaped where the work is judgment at the seams, deeply T-shaped where an operational failure has no acceptable blast radius. Organizations are notoriously bad at that kind of self-examination, which is exactly why it has to be decided rather than left to drift.</p>

<p>The observable that would change my mind is narrow and specific: if fluency and correctness converge — if the tools stop being confidently wrong — then the detection ability I have described has less work to do, and pure depth reasserts itself immediately. I do not expect that any time soon. But it is the thing I am watching, and if you want to argue with this essay, that is the productive place to do it.</p>

<h2 id="one-more-thing">One more thing</h2>

<p>The best technologists I have worked with in the last five years were all, without exception, M-shaped. None of them would have used the term, or comb-shaped either. They would have said they were curious, or interested in too many things. What they meant was that they could not sit still inside one domain, kept following problems across seams, and ended up with several depths because no single depth held their attention long enough.</p>

<p>Ten or twenty years ago that was an eccentricity, and it showed up on performance reviews as one.</p>

<p>Breadth is cheap now. Knowing which cheap answer is wrong is not — and nothing teaches that except having been deep, more than once.</p>]]></content><author><name>Antonio Elena</name></author><category term="engineering-org" /><category term="careers" /><category term="hiring" /><category term="ai-strategy" /><summary type="html"><![CDATA[Why pure specialization fails when AI makes breadth cheap — and why the naive generalist argument fails too. The second depth is what changes anything.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sig-intent.com/assets/img/og-default.png" /><media:content medium="image" url="https://sig-intent.com/assets/img/og-default.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">MD, JSON, HTML or XML: pick the format for the next reader</title><link href="https://sig-intent.com/writing/pick-the-format-for-the-next-reader/" rel="alternate" type="text/html" title="MD, JSON, HTML or XML: pick the format for the next reader" /><published>2026-05-10T00:00:00+00:00</published><updated>2026-05-10T00:00:00+00:00</updated><id>https://sig-intent.com/writing/pick-the-format-for-the-next-reader</id><content type="html" xml:base="https://sig-intent.com/writing/pick-the-format-for-the-next-reader/"><![CDATA[<p>It seems people are now starting to move to HTML for AI, instead of Markdown. Yes, certainly the informational density and the visual communication and design possibilities of HTML are far superior to what can be achieved with Markdown, and its mermaid or ascii diagrams, and code fences.</p>

<p>The argument — credited most cleanly to <a href="https://x.com/trq212/status/2052809885763747935">Thariq’s recent essay</a> (as of this writing) on Claude Code’s output formats on Twitter — runs roughly like this: HTML beats Markdown for AI-generated artifacts because HTML is both a data format and a rendering target. The same artifact serves machine parsing and human comprehension at once. Density compounds. The argument is correct. And I would add that it significantly raises the odds of people reading you, if you use HTML and not Markdown.</p>

<p>However, I think the debate and the framing is also incomplete in a way that matters. It is not HTML vs Markdown, rather, the question is: <strong>who is the next reader?</strong> Once you ask that, the format choice is not a debate, and becomes a matter of architecturally thinking downstream and making the best choice, something most teams never ponder.</p>

<p>This is a specific instance of not framing a decision at the right altitude. The format question is smaller than the <strong>decision-craft</strong> question beyond.</p>

<h2 id="what-the-format-debate-is-actually-about">What the format debate is actually about</h2>

<p>Let’s walk through the four formats in active use right now to find the right framing.</p>

<ul>
  <li>
    <p><strong>Markdown.</strong> The conversational format. Optimized for humans-and-LLMs reading prose with light structure: headings, lists, emphasis, code blocks. Cheap to produce, cheap to read, cheap for a model to attend to, portable. But it carries no metadata, no validation, no rendering richness — and that is the point. It is the format of the chat window, and perfect for that.</p>
  </li>
  <li>
    <p><strong>HTML.</strong> The battle-tested presentation classic format. Optimized for humans browsing a rendered artifact. Carries layout, navigation, tables, SVG, code, links, embedded interactivity and much more it accrued over the years. Density compounds because a single document serves both the parser and the eye. This is the format Thariq’s argument identifies and the reason Claude-generated dashboards and reports are increasingly delivered as HTML.</p>
  </li>
  <li>
    <p><strong>XML.</strong> Another veteran, the solid rugged contract format. Optimized for parsers, validators, schema-checkers, and language models trained to attend to tag boundaries. Its underrated feature is <em>mixed content</em> — the ability to put prose inside structure inside prose without escape-trapping it in a JSON string. Anthropic’s own prompting guidance leans on <code class="language-plaintext highlighter-rouge">&lt;context&gt;</code>, <code class="language-plaintext highlighter-rouge">&lt;example&gt;</code>, <code class="language-plaintext highlighter-rouge">&lt;instructions&gt;</code> for the same reason: the boundaries are explicit and the model treats them as load-bearing. Stricter and more verbose than the alternatives, and still a force to be reckoned with.</p>
  </li>
  <li>
    <p><strong>JSON.</strong> The control-plane format. Optimized for compact machine-to-machine traffic where payloads are small, fields are typed, and verbosity is overhead. Tool calls, API responses, configuration, agent control messages — JSON wins because nothing about the message is for a human, and brevity is a feature.</p>
  </li>
</ul>

<p>Each format has a primary downstream reader/consumer.</p>

<h2 id="the-reframe">The reframe</h2>

<blockquote>
  <p>Downstream thinking - Pick the format for the next reader.</p>
</blockquote>

<p>The consequence: most format arguments are not considering <em>who or what reads the artifact next</em>. Two engineers arguing whether agent output should be Markdown or XML are usually arguing whether the next reader is a human, another agent, a legacy system, or a downstream parser — and the disagreement is about that, not about syntax or verbosity. Resolve the consumer question and the format question becomes clear and takes about just a few minutes at most.</p>

<p>The corollary, which is where most teams actually go wrong: when an artifact has more than one next reader, <em>no single format is correct</em>. The honest answer is a layered one.</p>

<h2 id="the-hybrid-stack">The hybrid stack</h2>

<p>The pattern that pays off in practice, and which few think and build, is to choose the formats by the next reader-role rather than picking one one-size-fits-all standard or convention.</p>

<p>A specific example. An agentic system surfaces a finding from an incident-triage pipeline:</p>

<div class="language-xml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nt">&lt;finding</span> <span class="na">confidence=</span><span class="s">"0.7"</span> <span class="na">source=</span><span class="s">"trace-7421"</span><span class="nt">&gt;</span>
  <span class="nt">&lt;h3&gt;</span>P99 latency regression on /search<span class="nt">&lt;/h3&gt;</span>
  <span class="nt">&lt;p&gt;</span>Tail latency rose 40% after the 2026-04-12 deploy. Correlates with cache hit ratio falling below 85% on two of three regions.<span class="nt">&lt;/p&gt;</span>
  <span class="nt">&lt;evidence&gt;</span>
    <span class="nt">&lt;table&gt;</span>...<span class="nt">&lt;/table&gt;</span>
  <span class="nt">&lt;/evidence&gt;</span>
<span class="nt">&lt;/finding&gt;</span>
</code></pre></div></div>

<p>Three layers, one complete artifact with precise layered meanings, three readers:</p>

<ul>
  <li>The XML envelope (<code class="language-plaintext highlighter-rouge">&lt;finding&gt;</code>, <code class="language-plaintext highlighter-rouge">&lt;evidence&gt;</code>, attributes) is the <em>contract</em>. A downstream parser keys off it, a schema validates it, a confidence threshold filters on it before a human ever sees it.</li>
  <li>The HTML payload (<code class="language-plaintext highlighter-rouge">&lt;h3&gt;</code>, <code class="language-plaintext highlighter-rouge">&lt;p&gt;</code>, <code class="language-plaintext highlighter-rouge">&lt;table&gt;</code>) is the <em>interface</em> and the vehicle for visual communication. A reviewer renders it and reads it without translation.</li>
  <li>The attributes (<code class="language-plaintext highlighter-rouge">confidence</code>, <code class="language-plaintext highlighter-rouge">source</code>) are the metadata, structural and contextual.</li>
</ul>

<p>Generalize the pattern across a typical agentic system and the stack looks clear:</p>

<ul>
  <li><strong>JSON</strong>, or even yaml, for the control plane: tool calls, agent-to-agent commands, configuration, anything compact and typed.</li>
  <li><strong>XML</strong> for document envelopes and formal specs: schema-validated message structures, prompt blocks with explicit boundaries, intermediate representations between pipeline stages, payloads that need to round-trip between machines and humans without losing structure.</li>
  <li><strong>HTML</strong> for the visual and inspection layer: dashboards, audit logs, review screens, anything a human reads after the fact. Old technologies can be agently leveraged to transform the XML, not produced natively or generatively. That keeps the strictness as information moves.</li>
  <li><strong>Markdown</strong> for the conversation: model-to-human and human-to-model exchange, where light structure is enough and ceremony costs more than it pays. And now, also for human to agent communication, one order of magnitude richer.</li>
</ul>

<p>Thus:</p>
<ul>
  <li>XML for contracts</li>
  <li>HTML for visual interface and human communication</li>
  <li>Markdown for the conversation</li>
  <li>JSON is for the wire</li>
</ul>

<h2 id="the-cost-trap">The cost trap</h2>

<p>There is a discipline this pattern requires, and it is the part teams skip.</p>

<p><strong>Formalism only earns its keep when something downstream enforces it.</strong> XML tags, attributes, and schemas are not free. They impose verbosity, parsing cost, and a non-trivial cognitive load on whoever is writing prompts or reading raw output. The return is reduced ambiguity for parsers and models. If no parser, validator, or extraction step ever runs against the structure, the verbosity is dead weight.</p>

<p>The failure mode is predictable. A team adopts XML envelopes because Anthropic’s docs use them, or they believe in the value of <a href="https://aelena74.gumroad.com/l/xsp">XML Structured Prompting</a> in specific scenarios. They wrap every output in <code class="language-plaintext highlighter-rouge">&lt;response&gt;</code> and <code class="language-plaintext highlighter-rouge">&lt;reasoning&gt;</code> tags, but there is nothing downstream that parses the tags. The tags are thus just decoration, an illusion. They have paid the formalism cost without collecting the benefit, and the next engineer who joins assumes the structure is load-bearing and writes code around it that brittles fast from breaking changes upstream.</p>

<p>Pay for formalism only where it is enforced. Otherwise the simpler format is the correct call, and reaching for XML is over-engineering with extra steps.</p>

<h2 id="what-this-is-really-about">What this is really about</h2>

<p>The format debate is interesting because it is a clean instance of a much larger habit: arguing the answer before rushing to agree on the question. Most technology decisions get hastily resolved like that. Build vs. buy. Microservices vs. monolith. Synchronous vs. event-driven. SQL vs. NoSQL. The arguments stay heated because they are conducted at the wrong altitude — a layer below the reframe that would dissolve them.</p>

<p>The better move, in formats and in everything else, is one level of altitude up. Not the default choice or “which format do we use here?” but “who has to read this next, and how many of them are there?”. Not “build or buy?” but “is this differentiation surface or substrate?”. Not “monolith or microservices?” but “what is the unit of independent deployment our team can actually operate?”.</p>

<p>Get the altitude right and most arguments turn into routine engineering. Get it wrong and every meeting feels like litigation, because it is.</p>

<p>Remember: XML is for the contract. HTML is for the interface. Markdown is for the conversation. JSON is for the wire. Pick the format for the next reader, not the current writer.</p>]]></content><author><name>Antonio Elena</name></author><category term="ai-strategy" /><category term="architecture" /><category term="decision-craft" /><category term="agentic-systems" /><summary type="html"><![CDATA[Most format arguments are really arguments about who reads the artifact next. XML for contracts, HTML for the interface, Markdown for the conversation.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sig-intent.com/assets/img/og-default.png" /><media:content medium="image" url="https://sig-intent.com/assets/img/og-default.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">AI is an architecture problem, not just a model problem</title><link href="https://sig-intent.com/writing/ai-is-an-architecture-problem/" rel="alternate" type="text/html" title="AI is an architecture problem, not just a model problem" /><published>2026-05-09T00:00:00+00:00</published><updated>2026-05-09T00:00:00+00:00</updated><id>https://sig-intent.com/writing/ai-is-an-architecture-problem</id><content type="html" xml:base="https://sig-intent.com/writing/ai-is-an-architecture-problem/"><![CDATA[<p>The cheapest part of any AI feature is the model. The expensive parts are the ones nobody puts in the slide, and that get waved away in the discussion.</p>

<p>I have watched this pattern play out in three different use cases in the past eighteen months. A team picks a model, runs a POC in a week, and produces a demo that makes the room nod in vague agreement. Then the demo starts moving toward production and something quietly happens. A few  months later there is no product yet, no “value realization”. At best, there are unresolved tickets about data quality, retry semantics, evaluation pipelines, cost overruns, latency budgets, vendor lock-in, and a governance committee nobody invited. Most often, there is just silence.</p>

<p>The failure is actually on the system around the model.</p>

<p>Effectively shipping AI is a distributed-systems problem with ML as one component, and the gap between a working POC and a production product is almost never the model. It is the architecture and everything that has to be there to support the model.</p>

<h2 id="what-the-poc-does-not-show-you">What the POC does not show you</h2>

<p>When you run a POC, you get to assume the nice parts. The inputs are clean, the world is reduced in its messiness and complexity, there is no <a href="https://www.geeksforgeeks.org/system-design/fallacies-of-distributed-systems/">network fallacies</a>, no <a href="https://en.wikipedia.org/wiki/CAP_theorem">CAP theorem</a>, no latency, no bad data and no integrations or deployments. The call is synchronous. There is one user — you — and there are no failure modes to handle because the one call you made worked, so ship it because “it worked on my machine.”</p>

<p>Production is the very opposite of this set of assumptions.</p>

<p>Production has bad inputs, because real data is always worse in unpredictable ways. Production has concurrency, because you now have more than one user and some of them click twice. Production has latency budgets, because the user will abandon a feature that takes longer than their expectation of what it should cost. Production has a wide array of possible failures, because every call to a model is a network-mediated request to a stateful, throttled, regionally rate-limited third party that can and will degrade at the worst possible time. Production has a bill, because someone now adds up what each call costs and realizes the token budget and margin assumptions were wrong - or the compute, for that matter. And production has users who do adversarial things — intentionally or accidentally — because that is what users do with any interface that accepts free text.</p>

<p>None of those problems are model problems. All of them sit in the realm of distributed-systems problems. Every single one of them is well-known has been studied and solved in the context of service architecture for the past two decades, if not way more. The literature exists. Most AI teams are not using it because they think they are doing ML, and the ML literature does not cover it. Its focus is somewhere else.</p>

<h2 id="the-actual-components-of-a-working-ai-feature">The actual components of a working AI feature</h2>

<p>Strip away the model for a moment and look at what is actually required to run one in production. You need, at minimum:</p>

<p>A <strong>data pipeline</strong> that reliably delivers the inputs your model needs, with consistency guarantees that match the freshness needs of the feature. If you are doing retrieval-augmented generation, you need an index. If you are personalizing, you need user state. Both of those are real subsystems with their own failure modes. A data pipeline can easily be the hardest part to build, test, validate and maintain. And we know the quality of data in most organizations: not <a href="https://www.go-fair.org/fair-principles/">FAIR</a>, siloed, uncategorized, no MDM and so on.</p>

<p>A <strong>request pipeline</strong> that handles concurrency, rate limits, retries with backoff, and graceful degradation when the model is slow or unavailable, or puts limits on users’ actions. Every production LLM call should have a fallback — cached, cheaper, slower, or explicitly null. No fallback is a promise you cannot keep.</p>

<p>An <strong>evaluation harness</strong> that runs against a held-out set on every deployment and tells you whether your change made things worse. Drift hits from three directions: data drift as the inputs you retrieve and prompt with shift, model drift as your provider updates the endpoint underneath you, infra drift as the runtime around both moves with every deploy. Without this, you have no ability to iterate, you will be blissfully unaware shipping regressions you cannot diagnose.</p>

<p><strong>Observability</strong> that lets you answer “<em>what exactly did the model see, and what did it do?</em>” for any individual request a user complains about. You are logging the prompt, the context, the response, the latency, the cost, the user, and the outcome? Without this, how can you possibly triage and solve bugs?</p>

<p><strong>Cost control</strong> that tells you when a single user has consumed more of the month’s budget than you intended, and cuts them off or charges them. Without this, their consumption and the bill is your problem until addressed.</p>

<p>In many non-trivial environments, a <strong>governance layer</strong> that tracks what data went into the model, who authorized the integration, and what the model is allowed to do is a must-have. Skip this and your first regulatory question ends the product.</p>

<p>That is the whole system, and I am probably forgetting things. The model itself is but one box in it. If you cannot draw the system, you are not shipping AI; you are hosting a demo (in <code class="language-plaintext highlighter-rouge">localhost:3000</code> most likely).</p>

<h2 id="why-teams-skip-this">Why teams skip this</h2>

<p>Because it is boring, expensive and unglamorous. Decision makers often dismiss or underplay this complexity. A working production AI feature has maybe fifteen percent of its effort in the model and eighty-five percent in the system around it. The model is the part that demos; the system is the part that ships. Yes you might ship the demo faster, which lets you claim faster, which lets the team feel good in the short term. The dopamine hit.</p>

<p>Debt catches up with you in one of three ways.</p>

<p>First, cost — the bill arrives and someone realizes the unit economics do not work. Second, incident — the model has an off day, you have no fallback, the feature fails, and you cannot debug it because you logged nothing. Third, compliance — legal asks a question you cannot answer, and the feature is frozen until you can.</p>

<p>All three of these are predictable. All three are avoidable. None of them are solved by picking a better model.</p>

<h2 id="the-reframe">The reframe</h2>

<p>The thing to put on the wall when an AI project starts is not a model selection matrix. It is a system diagram. Every component above, every data flow, every failure mode, every cost lever. Do this <em>before</em> you pick a model, not after — because the model is the easiest thing to change and the hardest thing to get wrong in a way that kills the feature. Design for model swap-out from day one — ports and adapters, dependency injection, the strategy pattern. These predate the LLM era by twenty years and they remain the right tools for it.</p>

<p>What you will find, consistently, is that the model choice does not matter that much, or in some case, even barely matters. GPT-4, Claude, Llama, a hosted endpoint, a smaller fine-tuned model — all of them will work for 80%, if not more, of business use cases. What makes the difference between a stuck POC and a live feature is whether you have built a system that can operate one.</p>

<p>The companies that are shipping AI at scale right now are not the ones with the best models. They are the ones with the best architectures around the models. The model is a commodity. The system is the product.</p>

<h2 id="one-more-thing">One more thing</h2>

<p>If you are a CTO and your team is eight or whatever months into an AI initiative that is still at the POC stage, the question to ask is not <em>“are we using the right model?”</em>. The question is <em>“show me the system diagram that would need to exist for this feature to run in production tomorrow.”</em> If the team cannot draw one on a whiteboard in fifteen minutes, or point at one on the project wiki, the problem is not the model and it never was.</p>

<p>AI is an architecture problem. Treat it like one and you ship. Treat it like an ML problem and you will still be in POC next Christmas.</p>]]></content><author><name>Antonio Elena</name></author><category term="ai-strategy" /><category term="architecture" /><category term="distributed-systems" /><summary type="html"><![CDATA[Shipping AI in production is a distributed-systems problem with ML as one component. Why POCs stall at the demo-to-production cliff, and the reframe.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sig-intent.com/assets/img/og-default.png" /><media:content medium="image" url="https://sig-intent.com/assets/img/og-default.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>