How I Added Persistent Memory to a Competitive Intelligence Agent
How I Added Persistent Memory to a Competitive Intelligence Agent ๐ TL;DR: My CrewAI competitive intelligence pipeline forgot everything between runs. I added a persistent memory layer (Hindsight plus typed events), and it now spots patterns across weeks instead of summarising one week's news. I built a multi-agent pipeline that writes a cited competitive intelligence briefing. You give it a topic and some competitors, and CrewAI agents research the web, analyse what they find, and write a report. Every claim has to carry a citation that resolves to a real source. It worked, but only for a single run. Every run started blind. Ask it about a competitor in week 8 and it had no idea weeks 1 to 7 existed. It would summarise that week's news and call it intelligence. That's a lookup with good formatting. A real analyst sees five moves in a quarter and says "that's a pattern," and that conclusion needs history. So I built a persistent memory layer and put it in the middle of the pipeline. This post covers how it works, what changed, and what I got wrong. Before: four stateless agents The original crew was four agents in sequence: Discovery, Research, Analyst, Writer. Each passed its output to the next, and everything was thrown away at the end. The governance layer already worked, with a citation guard and a prompt-injection guard. The gap was time. After: a Memory Agent between Research and Analyst The pipeline now has seven agents: Discovery, Research, Memory, Analyst, Strategy Evolution, Prediction, and Writer. I use Hindsight as the persistent memory layer, with a class I wrote called HindsightStore in memory/hindsight_store.py acting as the application-level wrapper around it. My application still maintains structured event and profile data locally, while Hindsight provides the persistent memory and retrieval layer. The local structured data is used for deterministic profile calculations, strategies, and predictions.Each finding is stored as a typed event using a Pydantic schema, CompetitorEvent. It holds a competitor, an event type (feature launch, pricing change, hiring, acquisition, funding, partnership, market signal), a date, a title, a description, an impact score, a confidence value, and evidence URLs. I expose the memory functionality to my CrewAI agents through a HindsightStoreTool. The tool provides six operations: store_event, get_history, get_profile, search_memory, get_strategy, and get_predictions. The Memory Agent stores new findings through the memory layer and retrieves relevant historical context before the Analyst runs. It can use get_history and get_profile to combine persistent memory with the structured competitor profile, then passes the memory-enriched context to the Analyst. Strategy and Prediction agents can also retrieve relevant historical information before producing their outputs. Every write updates a derived profile Each store_event call also recomputes a CompetitorMemoryProfile for that competitor, with no batch job. Here's the core of it, from hindsight_store.py: hiring_events = [e for e in self._events if e.competitor == event.competitor and e.event_type == EventType.HIRING] if len(hiring_events) >= 4: prof.hiring_trend = HiringTrend.SURGING elif len(hiring_events) >= 2: prof.hiring_trend = HiringTrend.GROWING else: prof.hiring_trend = HiringTrend.STABLE high_impact = [e for e in self._events if e.competitor == event.competitor and e.impact_score >= 8.0] if len(high_impact) >= 3: prof.risk_level = RiskLevel.CRITICAL elif len(high_impact) >= 1: prof.risk_level = RiskLevel.HIGH elif prof.total_events >= 5: prof.risk_level = RiskLevel.MEDIUM prof.confidence_score = min(0.98, 0.3 + (prof.total_events * 0.07)) The confidence formula is the part that changes the output. It starts at 0.3, adds 0.07 per stored event, and caps at 0.98. The Writer's instructions say to hedge language for low-confidence competitors. A competitor I've seen twice gets cautious wording, and one with fifteen events gets firmer wording. Before vs. after I tested this on a fictional competitor, "NeuraCode AI". The store ships with seeded demo data, and scripts/demo_memory.py runs a store → recall → learn sequence with no LLM and no network. These are demo events, not real market data. Before: with only the latest event available, there is no historical context to compare against. The Analyst can only restate the latest item, something like "NeuraCode launched an AI code review tool." That's accurate but useless for strategy. After: with six stored events: feature_launch : AI code review tool launched hiring : 15 ML engineers hired pricing_change : Enterprise tier introduced at $45/seat/month acquisition : Analytics company acquired for $28M feature_launch : AI security scanner shipped partnership : Native JetBrains IDE integration announced The profile reports total_events = 6 and a confidence of 72% (0.3 + 6 × 0.07), straight from the formula. The Analyst now receives six dated events across product, talent, pricing, M&A, and distribution, and its task asks it to say where memory changed its conclusion compared with the latest news alone. This is a controlled demo, not a benchmark. I haven't run the pipeline on live competitors over weeks and measured briefing quality, so I'm not claiming a measured improvement. Why I Kept Typed Events Alongside Hindsight I kept CompetitorEvent as a structured application-level representation instead of putting everything into unstructured memory. This is useful when I need deterministic filtering such as: get_events(competitor="OpenAI", event_type="pricing_change", days=90) Hindsight handles persistent memory and retrieval, while my structured events support predictable date, event-type, impact, and profile calculations. The trade-off is that the structured layer is less flexible for semantic queries, while Hindsight can retrieve relevant historical context beyond exact event-type matching. The trade-off is that I lost semantic search. search_memory is a keyword scan over titles and descriptions. Search "cost reduction" and you won't find an event typed pricing_change unless those words are in its text. Guarding the memory Persistent memory can be poisoned, because a prompt injection that gets stored resurfaces in every later run. The injection guard strips instruction-like patterns from fetched pages before they reach the LLM. A separate memory_injection_guard() checks queries headed for the store, and validate_competitor_name() rejects names with shell metacharacters or over 200 characters. The citation guard runs on every briefing. What went wrong A bug in my own recency logic. The innovation score was documented as using a 90-day window, but it had no date filter, so old events inflated it forever. The code now has a comment marking the fix. Documenting a behaviour doesn't implement it, and only re-reading the code caught this. Impact scores are LLM guesses. The Memory Agent assigns each event a 0 to 10 score, and three events at 8.0 or above push a competitor to CRITICAL. That score isn't stable across models or small prompt changes, so risk levels and briefing text can shift with it. I want rule-based floors (large acquisitions floor at 8.0, for example) that the LLM can only adjust upward from. I haven't built that. Predictions never get graded automatically. update_prediction_status() exists, but nothing calls it in a loop. The system can't yet learn from its misses. Strategy parsing is brittle. crew.py parses the Strategy agent's output with regex, which breaks if the model varies its format. Schema-enforced output would fix that. A "fresh" store isn't empty. It auto-seeds demo data on first run, which is handy for demos and misleading for testing. ๐ What's next Based on the problems above, this is the order I plan to tackle them in: - Rule-based impact-score floors that the LLM can only raise - A loop that calls update_prediction_status() so predictions get graded - Schema-enforced output for the Strategy agent to replace the regex parsing - A run on live competitors over several weeks, so I can actually measure briefing quality ๐ Try it yourself Clone the repo below and run scripts/demo_memory.py. It needs no LLM and no network, so you can watch the store → recall → learn sequence in a few seconds. Takeaways - Store typed events, not summaries. You can't filter a summary by date or event type later. - Make memory use required in the task text. Optional memory gets ignored. - Tie output tone to evidence. Confidence that scales with stored events keeps the briefing honest. - Close the feedback loop early. Without prediction grading, memory only accumulates. ๐ Over to you: How are you giving your agents memory across runs? Typed events, vector search, or something else? I'd like to hear what's worked. Project Repository GitHub Repository Hindsight Resource HindSight Top comments (0)
Comments
No comments yet. Start the discussion.