Most Recruiter Bots Are Goldfish. I Gave Mine Hindsight.
The goldfish thing is a myth
Studies suggest real goldfish can remember things for months. Most recruiter bots I've seen can't manage a week. Here is what that looks like in practice: a candidate spends twenty minutes explaining what she wants, and three weeks later, the bot writes to her as if it has never heard of her.
It isn't broken. Each independent request only has access to the context we provide to it; without an external memory layer, earlier conversations aren't automatically available. That is the actual problem, and it is a plumbing problem, not an intelligence problem.
I built Recall to handle that plumbing: extract what a candidate said, store the useful parts somewhere durable, and retrieve the right context when a new role appears. Hindsight, the open-source agent memory layer from Vectorize, is what makes that loop possible.
The model is the goldfish, so wrap it
Once I stopped expecting the model to remember, the design got simpler. The model is a fast, stateless writer. Everything about what to say has to be fetched and put in front of it. Recall has four pieces:
- Express server - serves the interface and runs the workflow
- Groq (openai/gpt-oss-120b) - extracts facts from conversations and drafts recruiter-facing text
- Hindsight - provides long-term memory through
retain()andrecall() - Browser UI - shows extracted facts, recalled memory, and generated responses together
Groq doesn't store the candidate history. Hindsight doesn't decide what a recruiter should do. The application connects the two: the model extracts durable information, Hindsight stores and retrieves it, and the model uses that retrieved context to generate the next response. That separation also makes debugging easier. When something goes wrong, I can ask a concrete question: did the right memory come back, or did the model do something unexpected with the memory it received?
Step one: decide what is worth remembering
I don't retain raw conversations. Before anything is stored, the model reduces the conversation to short bullets covering things that may matter later: career goals, technical interests, location, work model, salary, work style, and constraints.
For Rashi Sharma, the first conversation produced information like:
- Career goal: backend engineering with more responsibility
- Technical interest: Node.js
- Location: Hyderabad first choice; open to relocation only for a highly suitable role
- Salary: target โน13-15 LPA; currently earning โน10 LPA
- Constraints: avoid frequent late-night work and high-pressure startups
Then those extracted facts are retained in Hindsight:
const importantFacts = response.choices[0].message.content;
await hindsight.retain(
BANK_ID,
`Important candidate information from a recruiter conversation with Rashi Sharma:\n\n${importantFacts}`
);
I could have used a rigid schema with fields such as salary_min and location. I didn't, because some of the most useful things candidates say don't fit neatly into database columns. "I'd rather not join another startup where weekends disappear" is a meaningful constraint, but forcing it into a predefined field can lose the actual meaning. Short extracted facts give the memory layer enough structure to retrieve useful context without pretending every candidate preference looks the same.
The important part is that I am not trying to make the model remember the conversation itself. I am creating a smaller, durable representation of what may matter later.
The test: one variable
The claim "memory improves the answer" is easy to fake, so I wanted a comparison that isolated the effect of memory. I kept the Groq model, candidate, opportunity, and question the same. The only difference is whether recalled Hindsight context is added to the prompt:
const context = withMemory && rememberedInformation
? `Recalled candidate context from Hindsight:\n${rememberedInformation}`
: "";
The interface makes this comparison explicit across Call 1, Call 2, and Call 3, with both "Without memory" and "With Hindsight" views.
Call 1 intentionally skips recall. The question is: How should I introduce this Backend Engineer opportunity to Rashi? Without previous candidate context, the response can accurately describe the opportunity:
Hi Rashi, we have a Backend Engineer role in Hyderabad (โน14 LPA, hybrid, Node.js, predictable hours). Would you be open to learning more about this opportunity?
It is accurate and polite. It would also work for almost any candidate with a similar technical background.
Later, Hindsight has Rashi's earlier preferences available. The same model can now produce a more specific follow-up:
Hi Rashi, just checking in-are you still prioritizing clear communication, reasonable working hours, and a strong work-life balance (and preferring to stay in Hyderabad without immediate relocation)?
The prompt template did not suddenly become smarter. The model did not change. The additional input is the candidate history retrieved from Hindsight. That is the part I care about. The difference isn't being attributed to a cleverer prompt or a different model. The memory is the variable.
The moment it matters
The practical problem behind all of this is the recruiter who asks, "Remind me, what's your salary range?" for the third time. With recalled context, the system can surface that Rashi currently earns โน10 LPA and previously stated a target of โน13-15 LPA. When a โน14 LPA role appears, the recruiter can immediately see that the listed compensation falls within the candidate's stated target.
But there is an important distinction between matching information and making a decision. Recall's recruiter brief has three parts: a summary, matches, and a verify list. The verify list is the part I would defend most strongly:
- Confirm the role does not involve a high-pressure startup environment or frequent overtime.
- Confirm the position offers the clear communication and backend ownership she previously said she wanted.
A job description saying "predictable hours" is a claim about the role. Rashi's preference for reasonable working hours is a fact about the candidate. The first does not prove the second will be satisfied. Recall therefore treats remembered preferences as context for the recruiter to verify, not as proof that the candidate should accept the role.
Asking for the right memory
Storing information is only half of the problem. Getting the right information back is where the design becomes interesting. I don't ask Hindsight for "everything about Rashi." Each workflow asks a question appropriate to the task:
- A follow-up can ask about goals, location, work model, compensation, and constraints.
- A new opportunity can ask specifically about compensation and preferences relevant to that role.
The retrieval call stays small:
const memory = await hindsight.recall(BANK_ID, query, {
includeSourceFacts: true,
maxSourceFactsTokens: 2000
});
The important design choice is the query, not just the fact that a memory store exists. The same candidate can have dozens of facts, but a recruiter asking about a particular Backend Engineer opportunity does not need all of them. Hindsight handles the retrieval layer, so I don't have to build a separate retrieval pipeline for every workflow. My responsibility is to ask a useful question and make the returned context available to the model.
That changed how I thought about agent memory. I stopped thinking of it as "load the user's profile" and started thinking of it as querying a history for evidence relevant to the current task.
The noisy bowl
The most useful bug I found wasn't a failed API call. It was duplicated memory. The interface displays the recalled memory alongside the generated brief. During testing, I noticed the same facts appearing repeatedly: Rashi's preference for clear communication and reasonable working hours showed up multiple times, and her salary was restated in several slightly different ways.
The memory wasn't wrong. I was. I had been retaining the same conversation on every test run. Hindsight stored each request faithfully, but repeated writes made the retrieved context noisier and harder to inspect. That led to a simple rule: retain once per conversation, not once per click.
It sounds obvious after seeing the problem. It wasn't obvious before I could see the accumulated memory. This is also why I think memory systems need to be observable. A UI badge saying "memory used" would have hidden the problem. Showing what was actually recalled made the duplication visible.
What I learned
- The model is stateless, so design around that. Don't ask the model to remember something it cannot access. Build the loop that retrieves the context it needs.
- Store what is worth keeping. Extracted facts are more useful here than blindly retaining every conversational turn, while still preserving preferences that don't fit rigid profile fields.
- Write recall queries for the task. A question about the current opportunity should return context relevant to that opportunity. "Tell me everything" is rarely a useful retrieval strategy.
- Test with one variable. Keeping the model, question, candidate, and opportunity fixed makes the effect of memory much easier to reason about.
- Make memory visible and keep the human in the loop. Showing recalled context makes the system debuggable, while the recruiter remains responsible for interpreting it and verifying anything that may have changed.
What I'd still fix
- Candidate data is personal. Compensation, work preferences, and reasons for leaving a job need authenticated access, per-candidate identity, consent, and mechanisms to correct or delete stored information.
- Retention also needs a clear idempotency and lifecycle strategy so the same conversation isn't written repeatedly and outdated preferences don't silently become permanent truth.
If I were building another agent that talks to the same person more than once, I'd design the memory loop before spending time tuning the prompt. The interesting part isn't making an agent say something clever. It is making sure that when the next conversation happens, the right piece of the previous one is still there.
Go deeper
For the underlying system, the Hindsight documentation for the retain and recall memory API explains how to get started, and Vectorize's overview of agent memory and persistent context provides the broader model for thinking about it. The implementation is available in the Recall recruiter memory agent repository.
Comments
No comments yet. Start the discussion.