You Aren't Choosing an AI Tool. You're Choosing Who Gets Paged at 2 AM.
DEV Community

You Aren't Choosing an AI Tool. You're Choosing Who Gets Paged at 2 AM.

Part of AI Leadership in the Real World - how leaders turn scattered pilots into governed, adopted, measurable capability. TLDR: A support agent doing 50,000 chats a month needs ~3.5 FTE and $500k+/year just to stay accurate - while a typical 100-seat Copilot rollout sees only 20-30 seats used weekly. For SMBs, build-vs-buy isn't about features. It's about what you can afford to own for 24 months. We thought we were choosing a tool. We were really choosing a future dependency, a support queue, a governance burden, and a second bill that arrives a year later. Every vendor demo promised acceleration, control, and simplicity at once. Every internal proposal promised flexibility, ownership, and leverage. Nobody said both bills arrive late - one in engineering on-call, the other in consumption meters. Good platform decisions feel a little boring at first and very smart a year later. Why AI is special (and why old build-vs-buy math breaks) Traditional software mostly stays still when you leave it alone. AI doesn't: It drifts. Knowledge changes, customer language shifts, users ask harder questions once they trust it. Accuracy quietly drops from 90% to 70% with no error log. It speaks for you - legally. A wrong Confluence page is embarrassing. A wrong chatbot answer is a commitment a tribunal can enforce. It lives on someone else's deprecation clock. OpenAI gives at least 6 months before retiring a GA model. That's a hard deadline, not a backlog item. Prompts, evals, and output parsers all need rework. It multiplies cost per request. One human click = one action. One agent resolution = 6 lookups, drafts, updates, and logs - each potentially metered. It turns connectors into permanent work. Salesforce, SharePoint, Jira, Zendesk all change auth, rate limits, and APIs. Your agent keeps running while its knowledge goes stale. Gartner predicts 40%+ of agentic AI projects will be canceled by end of 2027 on cost, unclear value, and weak risk controls. McKinsey's State of AI 2025 (Nov 5, 2025) found 88% report regular AI use in at least one function, yet only 39% report any enterprise EBIT impact (most under 5%; ~6% high performers). BCG's Widening AI Value Gap (2025, n=1,250) found only 5% generating value at scale while 60% see minimal gains. The gap isn't building. It's owning. Here are 6 public cases every SMB CTO should know - scenario, what happened, what didn't work, and why. 1. The chatbot that created legal liability: Air Canada (2024) Scenario: Jake Moffatt, booking emergency travel after his grandmother died, asks Air Canada's website chatbot about bereavement fares. The bot says: book full fare now, claim the discount within 90 days. What happened: That was wrong. The real policy barred retroactive claims. Moffatt flew, applied with screenshots and a death certificate, was denied. He took it to the British Columbia Civil Resolution Tribunal - Moffatt v. Air Canada, 2024 BCCRT 149 (Feb 14, 2024). The tribunal ordered Air Canada to pay C$812.02 (C$650.88 fare difference + interest + fees) for negligent misrepresentation. It explicitly rejected Air Canada's "remarkable" defense that the chatbot was a separate entity responsible for its own actions: "it is still just a part of Air Canada's website." What didn't work and why: No grounding to the actual policy page, no policy-layer guardrail, no correction path. The bot hallucinated a generous policy linked to the real bereavement page that contradicted it. SMB lesson: If you ship a customer-facing agent - bought or built - you own what it says. For a 20-person company, one such incident is a support crisis and a trust crisis. Budget the eval + human escalation before launch, not after. 2. The "AI platform" that was mostly services: Builder.ai (2025) Scenario: Builder.ai, founded 2016 as Engineer.ai, marketed "Natasha" as AI that builds software "as easy as ordering pizza." Raised $450M+, valued at ~$1.5B in 2023 with Microsoft and Qatar Investment Authority backing. What happened: On May 20, 2025 it entered insolvency. Audits revised claimed 2024 revenue of $220M down to ~$55M. Creditor Viola Credit seized ~$37-40M after a $50M debt facility. The US Attorney's Office for SDNY subpoenaed records. UK entity Engineer.ai Global Ltd went into compulsory liquidation. Viral coverage said "700 engineers faked the AI." That shorthand is contested - Gergely Orosz (Pragmatic Engineer), after talking to ex-employees, corrected it: there was a real AI team doing spec, estimation, and code workflows, alongside hundreds of outsourced developers doing delivery. WSJ had flagged heavy human reliance as early as 2019. What didn't work and why: The hybrid model can be legitimate. The failure was marketing opacity + financial misrepresentation: buyers couldn't tell which part was automated, which was human, and what would happen if the vendor collapsed. Customers who outsourced their roadmap to it lost the platform overnight. SMB lesson: When buying an "AI platform," diligence the labor model: what % is model vs. reusable components vs. humans? What happens to your code, data, and uptime if they go under? Get escrow and export in writing. 3. The seat license nobody sat in: Microsoft 365 Copilot Scenario: A 100-person SMB buys 100 Copilot for Microsoft 365 seats at $30/user/month ($36,000/year) - plus the required underlying M365 license upgrade many teams forget to price. List price: Microsoft lists Microsoft 365 Copilot at $30/user/month (annual commitment) as an add-on requiring a qualifying M365 base plan - true all-in $42-$90/user/month depending on base tier. Copilot Studio is separate consumption: ~$200/mo per 25,000-credit pack. What breaks: Independent utilization surveys consistently report only 20-30% of purchased seats see weekly active use at scale (directional, not audited). Change management (training, comms, support) is repeatedly estimated at 30-50% of license cost and rarely in the proposal. Separate cautionary tale (consumer, not enterprise): Australia's ACCC filed Federal Court proceedings Oct 27, 2025 alleging Microsoft misled ~2.7M Personal/Family subscribers by bundling Copilot rises ($109โ†’$159 Personal, $139โ†’$179 Family) while hiding the cheaper Classic no-AI option in the cancellation flow. The widely cited 100%+ ROI figures trace to a Microsoft-commissioned Forrester TEI study - real data, not independent data. SMB scale-down: At 25 seats instead of 100, the waste is smaller in dollars but identical in ratio - 5-8 active seats carry the other 17. That's why seats must be earned, not allocated. What didn't work and why: Per-seat pricing assumes uniform adoption. In SMBs adoption is spiky: 15 power users love it, 60 try twice, 25 never enable it. You pay for 100, get value from 25. The prerequisite upgrade + idle seats kills ROI by month four. SMB lesson that worked: Pilot with a 25-seat power cohort, instrument weekly active use from day one, and only expand when active-use >60% for 4 weeks. Disciplined teams treat seats as earned, not allocated. 4. The second meter: Salesforce Agentforce ($2/conversation) Scenario: A mid-size team handling 50,000 service interactions a month turns on Agentforce. List price: $2 per conversation - on top of existing Service Cloud seats. List price: $2 per conversation at list, before discounts - and before the Service Cloud seats underneath. Salesforce's current official model is Flex Credits ($500/100k; 1 standard action = 20 credits = $0.10) as the flexible alternative, with unused credits not rolling over and Flex vs. Conversations not mixable in one org. What breaks: Conversation definitions (failed/escalated/abandoned handling) and per-workflow action counts are contract-specific - which is exactly why pre-sign modeling matters. Internal IT/HR agents with lots of short queries get the worst unit economics on per-conversation pricing - same $2 whether the answer saved $200 or $2. SMB scale-down: At 5,000 interactions/mo instead of 50,000, the meter drops 10x (~$10k/mo at list) - but the modeling work doesn't. You still must define the billable unit in writing before signing. What didn't work and why - usage amplification: A human resolving a ticket = 1 action. An agent resolving it = query record + search KB + draft response + update case + send email + log interaction = 6 metered steps, depending on config. Teams estimated on user requests but were billed on agent actions. Copilot Studio has the same pattern: $200/mo for 25,000 credits, but generative actions burn credits faster than classic flows, plus separate Azure token + compute charges. Three bills, three consoles. SMB lesson that worked: Demand a cost-modeling pilot in writing: define "conversation/credit" for your workflow, exclude tests, cap overages, negotiate rollover + 24-month price lock. Design agents to cache lookups and confirm intent before chaining actions - teams report 20-40% consumption cuts from flow design alone. 5. The build that became a hidden product: the 3.5-FTE support agent Scenario (composite SMB, modeled on SearchUnify's 50k/mo reference - not a single named company): A 30-person SaaS company builds a RAG support agent on LangGraph + Confluence + Zendesk. Pilot hits 88% answer accuracy. CEO calls it "done." What happened (per SearchUnify's 2026 vendor field analysis - directional, not audited): To keep a 50k-interaction/mo agent production-grade you need roughly: 1.0 AI/ML engineer + 1.0 data engineer + 0.5 platform + 0.5 security/governance + 0.5 product analyst = 3.5 FTE, ~$500k-700k/year before infra, tokens, observability, and audits. Major model migrations take 6-10 weeks each (re-benchmark, re-prompt, re-test, re-secure). Connectors drift: auth changes, endpoints retire, docs grow by thousands of pages. Three drifts compound: knowledge drift (docs change), data drift (tickets change), behavioral drift (users ask harder questions). Without evals, you find out from angry customers. SMB scale-down: At 5k interactions/mo the token/infra meter d

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.