The API Gateway Shouldn't Know Your Business
DEV Community

The API Gateway Shouldn't Know Your Business

I appreciate what an API gateway brings to most systems, at the right time and in the right place. It gives every service a shared first layer. TLS, rate limiting, and access logging happen once at the edge, and bad tokens get turned away before any service spends work on them. The services behind it still check what reaches them, but they start from cleaner traffic. The gateway also lets those services change without breaking their clients. But a gateway's configuration tends to collect things nobody designed it to hold, like a check on a customer's plan, a response field renamed back for an older client version, or a call to a second service to decide whether the first one should be reached at all. That drift, away from the gateway's purpose and away from the code that owns the rule, isn't always carelessness. The gateway is the cheapest place in the system to apply a business domain rule to every external request, so rules flow toward it. The decade-old advice that business logic doesn't belong there hasn't stopped the drift for many organizations, because it says where business logic shouldn't go and little about the forces that send it there, or about how to keep a rule out when the need is urgent. A business domain rule belongs in the gateway only if the team that owns the gateway could change it correctly without asking a domain team. Keeping the rest out takes an answer to each force, and an emergency rule that does go in goes in as debt, with a domain owner and a scheduled removal. The Advice Has Been Clear Since 2014 In March 2014, James Lewis and Martin Fowler described microservices as favoring "smart endpoints and dumb pipes," against the Enterprise Service Bus, whose products "often include sophisticated facilities for message routing, choreography, transformation, and applying business rules." Thoughtworks put "overambitious API gateways" on Hold in its Technology Radar from November 2015 through May 2018. The April 2016 entry calls the pattern "a worrying re-emergence of this disease," the disease being business smarts pushed into middleware, and says that "any domain smarts such as data transformation or rule processing should live in applications or services where they can be controlled by product teams working closely with the domains they support." Vendors' own guidance agrees. Microsoft's Gateway Offloading pattern, the Azure Architecture Center's guidance on moving shared concerns like TLS termination, authentication, and throttling into a gateway, says "Never offload business logic to the gateway." AWS's documentation for REST API mapping templates recommends a proxy integration over transforming data in the gateway when possible. Why Rules Drift Into the Gateway Four forces pull a rule toward the gateway, and each one makes sense to whoever is making the change that day. Every External Request Passes Through It Suppose a rule has to apply to every client of an API. Premium customers can request to export more than 10,000 rows, and everyone else can't. In the services, that rule is a change to the export service, and possibly to each service that exposes a large download. In the gateway, it's one policy that looks up the caller's plan, compares it with the row count in the request, and rejects the request before any service sees it. The person making the change may not even see it as a pricing rule, because in the moment it can look like an urgent fix to stop large exports from degrading the system for everyone. It Ships on a Different Schedule A gateway policy usually ships as configuration, and a platform team often owns it, so adding the export limit there doesn't wait on the export team's backlog. What the gateway offers is a different queue, not necessarily a faster one. When the export team is busy and the deadline is close, the rule goes wherever it can ship this week. After that, every change to the export limit waits in the platform team's queue, even when the export team could have shipped it sooner. The Gateway Offloading pattern says offloading doesn't fit when centralizing concerns "creates a change-management bottleneck" because the gateway team's release cycle is slower than the service teams'. The schedule that let the rule in once is the schedule it's stuck with. Vendors Compete on Programmability A rule can only move into the gateway if the gateway can express it, and gateway products have made sure it can. Azure API Management runs policy expressions written in C#, with conditional choose blocks and a send-request policy that calls another service and stores its response for later policies to read. Kong's Pre-Function plugin "lets you dynamically run Lua code" inside the gateway. Thoughtworks' later radar entries traced the trend to vendors in a highly competitive market adding features to differentiate their products. Suppose the export team is booked for the quarter, so leadership overrides the queue and hands the export limit to the platform team, which runs Azure API Management. The platform team writes a policy that takes four steps on each export request: - Read the customer ID from the caller's validated token. - Ask the entitlements service for that customer's plan with send-request , caching the answer for a few minutes so each export doesn't pay for a second call. - Read the requested row count from the rows query parameter, treating a missing value as zero. - Return a 403 if the plan isn't premium and the count is over 10,000. Every step uses a standard gateway feature, and the result is business logic. Once the plan lookup exists, the next rule that needs a customer's plan costs one more condition. Security Wants One Place to Check Authorization is often the first business rule to land in the gateway, pushed by a security group that wants one place to audit rather than trusting each dev team to get it right. The gateway already checks tokens, so it looks like that place. The drift starts when the gateway moves from checking that a token is valid to deciding what the caller may do, usually with a role check on a route that grows into decisions about which records the caller may touch. A Gateway Rule Loses Its Owner, Its Coverage, and Its Tests In the owning service, a rule has one team deciding what it means, runs on every path to the capability it protects, and is tested with the code it governs. In the gateway, it can lose all three. Its Meaning and Its Enforcement Get Different Owners Go back to the platform team's export policy, which refuses an export over 10,000 rows unless the customer is on the premium plan. That one condition encodes two business facts, that a plan called premium exists and that it allows larger exports than every other plan. The entitlements team decides what plans exist and the export team decides what they allow, but the platform team owns the policy. The Gateway Offloading pattern presents a dedicated gateway team as a benefit for specialized concerns like security, but for a business rule it means the team that can change the rule's meaning can't see where it's enforced, and the team that enforces it can't judge whether it's still right. Consider what happens when sales introduces an enterprise plan. The entitlements service starts returning enterprise , and the gateway rule still compares against premium , so enterprise customers, who pay the most, get rejected on large exports. No test fails, because no test in either service ever ran the rule. The fix needs the platform team to change a rule whose meaning they didn't write, once someone actually realizes it exists. The same split makes the rule hard to remove. The platform team can see the rule but can't tell whether any client depends on it, and the domain team could answer that but may not know the rule exists. So rules accumulate, each one cheaper to leave than to trace, until the gateway becomes the middleware Lewis and Fowler described. The split happens only when separate teams own the gateway and the service. A team that owns both keeps one owner but still loses paths and tests, because those losses come from where the rule runs. The Rule Covers Only What the Gateway Can See The gateway can enforce the rule only on requests that pass through it, and only with what those requests contain. The first gap is the routes that never cross the gateway. Follow one customer on the free plan who wants 50,000 rows. Calling the public API, they get the 403. Then they click Export in the product's web app, whose backend calls the export service over the internal network, and the export runs. They set up a nightly scheduled export. The scheduler, acting with its own service identity, publishes an ExportRequested message that the export service consumes, and that export runs too. The same customer asks the same question three ways and gets two different answers. The Gateway Offloading pattern recommends that backends accept requests only through the gateway, which stops outside clients from bypassing it but doesn't help here. The web app's backend and the scheduler are part of the product, the queued message never becomes an HTTP request, and even a scheduler call routed through the gateway would carry the scheduler's identity, with no customer for the gateway to look up. So the export service needs its own check, and the enterprise-plan change now has to reach two copies of the rule. Without that check, free customers get unlimited exports by scheduling them. The second gap is on the gateway's own route. The policy reads rows from the query string, because a declared row count is all the gateway can see before the export runs. A request that leaves rows out defaults to 0, and one that asks for ?from=2020-01-01&to=2026-01-01 never mentions rows at all. Both pass, and the service returns however many rows the query matches. Only the service runs the query, so only the service can enforce the limit on what an export actually produces. The Rule Leaves the Domain's Tests The export team

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.