Why I can't run Splink in an API call
DEV Community

Why I can't run Splink in an API call

The Latency Problem

The 2asy.ai news pipeline uses Splink. Company mentions come in with each article batch, Splink runs pairwise comparison against everything already in the registry, clusters form, the registry updates. Usually takes about 90 seconds. Sometimes longer. Nobody is waiting for it.

When I started building the ER API, I thought the API would just be a lighter version of that. Call Splink, get a match, return it. The response latency problem hit me within the first day of prototyping.

Splink's actual comparison work involves blocking across the full corpus, scoring candidate pairs, then an EM step to calibrate probabilities. On a corpus with a few hundred thousand records, you're looking at seconds of compute, not milliseconds. You can't put that in the path of an HTTP request that someone's polling from their application.

The Workaround and What It Costs

So the API compares an incoming name against a pre-computed candidate index and returns the top result above a threshold. That's fast enough.

What I didn't think through carefully at the start was what you lose by doing it that way. Splink's EM step sees the full distribution of records together. Some pairs that look ambiguous in isolation are only resolvable in context because the EM algorithm essentially gets to vote across the whole population. The API doesn't have that. Each query is independent.

The Threshold Question

I worked around the latency problem reasonably well. The threshold question has been messier.

When a match comes back at 0.71, what should the caller do with that? In the news pipeline I can defer uncertain pairs to a review queue and come back to them. In the API you have to give the caller something.

My first attempt was to just return the score and let callers use their own threshold. What I didn't anticipate was that most callers treat any result above their threshold as a hard match, regardless of how close to the threshold it is. The docs now include threshold guidance for different use cases, but I'm still getting questions about this, so I probably haven't described it clearly enough yet.

Model Retraining

Model retraining is the one I thought about least and arguably got most wrong. I collect labeled pairs continuously and retrain the Splink model for the pipeline every few weeks. Doing the same for the API turned out to be a problem.

Even a small retrain shifts match probabilities slightly. If the API returns 0.84 for a query this week and 0.79 for the same query after a retrain, anyone who built automation against the API and didn't read the changelog will notice. I now pin the API to a specific model version and do explicit releases instead. This is more conservative than I'd like, but I don't have a cleaner answer.

Different Use Cases

The 2asy.ai pipeline still uses full Splink batch resolution. The API handles point queries where a caller needs an answer immediately and doesn't have the corpus context to run it themselves. I knew from the start that these were different use cases. What I underestimated was how different the constraints would be.

I build er-api, a multilingual entity resolution service for Korean, Japanese, Chinese, and English corporate data. More at hannune.ai.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.