In 2015 I ran the platform evaluation for a rebuild.
I dismissed the option that would go on to define the next decade of my work.
My assessment was that it was a Lego-block platform. Fine for small internal tools. Not capable of carrying the data volume and operational complexity the business needed.
I was wrong.
The interesting part is not that I got it wrong. Evaluations get things wrong. The interesting part is why, because the same error is being repeated right now by people choosing AI platforms.
I was evaluating the platform against the code we would write.
The question was the delivery system we would have to run.
Disclosure: I now work for a partner of the platform I originally dismissed. That relationship is relevant to how you read this. It is also why the part worth writing down is the reasoning error rather than the product.
The evaluation I actually ran
The shortlist was reasonable: the platform in question, Mendix, a few other vendors, and the option of building our own stack and software development lifecycle from scratch.
The way I compared them was not reasonable.
I tested what a developer could express. How much control there was over the generated output. Whether an unusual requirement could be satisfied without fighting the tool. Where the ceiling was on customization.
Every one of those questions is about authorship.
None of them is about operation.
We were a small team inside a business that already had customers, already had regulatory obligations, and could not stop serving either during a rebuild. That is the actual condition described in modernization is a capacity decision before it is a technology decision. The constraint was never expressiveness. It was how much delivery capacity we had, and how much of it the platform would consume or return.
Building our own stack scored well on authorship. It would have been a catastrophe on capacity.
What changed the decision was not a demo
This is where I have to be careful, because I have argued the opposite case.
In a conference demo is not a production decision I made the point that a prepared stage removes the variables production puts back, and that a demo is not evidence of readiness. I still believe that.
What changed my mind was at a conference. It was not a demo.
An insurance company from Europe presented their work. They were not selling anything. They described the data volume they were running, the complexity of the domain they operated in, and what it took to keep it working.
That is a different evidence class entirely.
A vendor demo shows a capability under conditions the vendor chose. An operator account shows a system under conditions the operator did not choose. One is a claim about what is possible. The other is a report from inside the constraint.
I sat there doing the arithmetic. If they could hold that volume and that complexity, our load was not the problem I had assumed it was.
The ceiling I had imagined was not where I thought it was. I had inferred it from the surface of the tool rather than measuring it against anyone actually running at scale.
| Evidence | What it demonstrates | What it cannot tell you |
|---|---|---|
| Vendor demo | A capability exists under prepared conditions | Whether it survives your constraints |
| Reference call arranged by the vendor | Someone succeeded, selected for having succeeded | What the failure distribution looks like |
| Practitioner account under load | A real operating envelope, including the parts that hurt | Whether your domain is comparable |
| Your own pilot | Behavior in your environment at small scale | Whether the organization can own it |
None of these is sufficient alone. The mistake I made was skipping the third row entirely and substituting my own impression of the tool.
Toolkit and delivery system are different purchases
The specific thing I had misclassified was the category.
I had it filed as a low-code tool. A faster way to produce an application. Something you use to write software.
It was a full software development lifecycle. Source control, environments, staged deployment, dependency analysis, impact tracking, and the path from a change to production, all governed as one system.
Those are not the same purchase.
A toolkit changes how a developer spends an afternoon.
A delivery system changes what an organization can promise, how quickly it can respond, what happens when a person leaves, and whether a regulator or an acquirer can be shown how a change reached production.
If you evaluate a delivery system as though it were a toolkit, it will look constrained. That is the correct reading of the wrong object. The constraints you are seeing are the governance, and the governance is the product.
I saw the constraints. I did not understand that I was looking at the feature.
What the evaluation should have asked
The questions that would have produced the right answer in 2015 have nothing to do with syntax.
- How long from a decision to that change being live for customers, including review and approval.
- What happens to delivery velocity when the team doubles, and when it halves.
- How long before a new developer is productive without supervision.
- What the upgrade path costs, and who absorbs it.
- What evidence the system can produce about how a change reached production.
- What has to be true for this to survive technical due diligence.
- Where the work goes when the platform cannot do something.
That last one is the honest test. Every platform has a boundary. The question is whether the escape route is designed or improvised.
I now think the last two matter more than anything on my original list. We went through a diligence process in 2019, and the ability to show how software reached production was worth considerably more than any amount of expressive control would have been.
Text version
| Toolkit evaluation asks | Delivery-system evaluation asks |
|---|---|
| What can a developer express? | How does a change reach production? |
| How much customization is possible? | What happens when the team grows or shrinks? |
| Can unusual requirements be coded? | How quickly can a new developer become productive? |
| How much control exists over generated output? | What does upgrading cost and who absorbs it? |
| Where is the technical ceiling? | What evidence exists for audit and diligence? |
| How do we escape the platform? | Is the escape route designed or improvised? |
The same error, on a new platform
This is not a story about one product decision in 2015. I would not write it up for that.
I am watching the identical reasoning error happen with AI platforms right now.
Teams evaluate a model or an agent framework by what it can produce. Quality of output. Flexibility of prompting. Whether an unusual case can be handled. All authorship questions, exactly the ones I asked.
The delivery questions are the ones that decide the outcome. Who approves an agent’s action. How a change to a prompt or a tool definition reaches production. What evidence exists when something goes wrong. What happens at the boundary of what the system can do. Whether the organization can own the thing after the team that built it moves on.
This is the same gap between possibility and ownership behind the pilot is not the product, and the same reason a company cannot scale what it has not made repeatable.
A capability you cannot govern is not an asset. It is a dependency you have not priced.
What I take from being wrong
Three things.
The first is that a platform evaluation is an operating decision wearing technical clothing. If the evaluation only involves engineers assessing expressiveness, the wrong question is being asked well.
The second is that practitioner evidence outranks vendor evidence, and it is not close. The most useful twenty minutes in that entire evaluation was someone with no stake in my decision describing what their system actually did under load. Vendors will arrange reference calls. Those are selected. Go find the operators nobody selected.
The third is the one that took longest to accept. My judgment was confident, experienced, and wrong, and nothing inside my own reasoning would have corrected it. It took contact with someone else’s production reality.
That is worth remembering the next time an evaluation feels obvious.
The platform I dismissed carried a rebuild, a client base that grew from 44 to more than 300, a migration of more than 50TB with no downtime, and a technical due diligence process. I have been building on it since.
I was not wrong about what I saw.
I was wrong about what I was looking at.

