Building a Ride-Share Zone-Balancing Agent with LangGraph - Part 3: Giving the Agent Memory
This is Part 3 of a 5-part series. Part 1 built a rule-based agent for one zone. Part 2 added an LLM for two narrow jobs: reading an ops note, and explaining the decision. Neither part changed one basic fact: the agent looks at a zone exactly once, then stops. That's a real problem. A severe deficit often needs more than one 15-minute window to fix. But calling the agent again is first contact all over again: - No memory of what was already tried - No sense of whether this is cycle one or cycle six - The zone's own numbers reset, instead of carrying forward Part 3 fixes that with MemorySaver and a thread_id . One thread becomes one zone's ongoing story. One .invoke() call becomes one 15-minute cycle. The decision logic itself doesn't change at all. detect_imbalance , classify_severity , resolved_imbalance , and choose_best_policy are the exact same deterministic functions from Part 1, reused directly. What memory actually buys here: - The zone's real state carries forward correctly, instead of resetting. That includes drivers, riders, and a little organic replenishment each cycle. - A note can schedule something for future cycles - "traffic bad for the next hour," "the event starts at 6pm" - without being repeated every time. The LLM extracts the fact once. Code applies it automatically from then on. The Graph Two new nodes get added at the front. Both are plain Python - no LLM: start_cycle โโโถ apply_scheduled_conditions โโโถ detect_imbalance โโโถ classify_severity โโโถ set_candidates โ โโโโโโโโโโโโโโดโโโโโโโโโโโโโ โผ (balanced) โผ (deficit/surplus) trivial_do_nothing reconcile_inputs โ LLM #1 โ โผ โ resolved_imbalance โ โผ โ choose_best_policy โ โผ โ generate_explanation โ LLM #2 โ โ โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโ โผ simulate_and_report - start_cycle clears the previous cycle's raw conversation. This keeps the message log from growing forever across separate.invoke() calls. (History still persists - it's kept in its ownhistory field, not in the message log.) It also advances the clock, so demand and supply use the right hour's baseline. - apply_scheduled_conditions checks whether anything a past note scheduled should expire or trigger this cycle. No LLM call - just comparing the current cycle number against a stored expiry or trigger. reconcile_inputs also gets extended, not replaced. It's the same single-shot call from Part 2. Now it also picks up a stated duration ("for the next hour") or a future trigger time ("starts at 6pm"), when the note mentions one. The LLM extracts the raw fact once. apply_scheduled_conditions applies it automatically on every cycle after that. The State AgentState grows four fields past Part 2's: - initial_hour andcycle_number anchor the clock - history accumulates one record per cycle - scheduled_conditions holds whatever a past note scheduled for later None of this is MemorySaver itself - it's just ordinary state. MemorySaver is what makes this state survive between separate calls, instead of resetting every time: class AgentState(TypedDict): zone: dict ops_note: str initial_hour: int # set once on cycle 1, anchors the clock cycle_number: int history: Annotated[list[dict], operator.add] scheduled_conditions: dict # e.g. {"traffic": {"expires_at_cycle": 5}} imbalance_ratio: float imbalance_type: str severity: str candidate_policies: list policy_evaluations: dict policy_resolutions: dict recommended_policy: str explanation: str messages: Annotated[list[BaseMessage], add_messages] outcome: dict outcome_delta: dict report: str The Extended Tool report_context_and_schedule is Part 2's report_context tool, with two fields added. That lets the LLM report timing, not just a current condition - a stated duration, or a future trigger hour. Code still owns all the arithmetic that follows: minutes to cycles, hour to cycles-until. The LLM only extracts the raw fact: @tool def report_context_and_schedule( rain_flag: bool, event_flag: bool, traffic_level: Literal["none", "light", "moderate", "heavy"], traffic_duration_minutes: int = 0, event_starts_at_hour: int = -1, ) -> str: """ Call this exactly once with your interpretation of the ops note for THIS cycle. If the note doesn't affect a value, report the zone's current value unchanged. traffic_duration_minutes: ONLY if the note states how long this traffic condition will last (e.g. "for the next hour" -> 60). 0 if no duration is stated -- traffic_level then applies to this cycle only. event_starts_at_hour: ONLY if the note describes an event scheduled for a specific clock time that hasn't started yet (e.g. "starts at 6pm" -> 18). -1 if no future start time is stated. """ return ( f"rain_flag={rain_flag}, event_flag={event_flag}, traffic_level={traffic_level}, " f"traffic_duration_minutes={traffic_duration_minutes}, event_starts_at_hour={event_starts_at_hour}" ) Wiring In the Checkpointer Every node here is a plain function, same as before. The one genuinely new piece is MemorySaver , passed in at compile() time. That's what turns zone /history /scheduled_conditions from ordinary state into something that survives across separate .invoke() calls on the same thread_id : llm = ChatOllama(model="qwen2.5:14b", temperature=0) llm_with_reconcile_tool = llm.bind_tools([report_context_and_schedule]) def _reconcile(state): return reconcile_inputs(state, llm_with_reconcile_tool) def _explain(state): return generate_explanation(state, llm) g = StateGraph(AgentState) g.add_node("start_cycle", start_cycle) g.add_node("apply_scheduled_conditions", apply_scheduled_conditions) g.add_node("detect_imbalance", detect_imbalance) g.add_node("classify_severity", classify_severity) g.add_node("set_candidates", set_candidates) g.add_node("trivial_do_nothing", trivial_do_nothing) g.add_node("reconcile_inputs", _reconcile) g.add_node("resolved_imbalance", resolved_imbalance) g.add_node("choose_best_policy", choose_best_policy) g.add_node("generate_explanation", _explain) g.add_node("simulate_and_report", simulate_and_report) g.add_edge(START, "start_cycle") g.add_edge("start_cycle", "apply_scheduled_conditions") g.add_edge("apply_scheduled_conditions", "detect_imbalance") g.add_edge("detect_imbalance", "classify_severity") g.add_edge("classify_severity", "set_candidates") g.add_conditional_edges("set_candidates", route_llm_or_skip, { "trivial_do_nothing": "trivial_do_nothing", "reconcile_inputs": "reconcile_inputs", }) g.add_edge("reconcile_inputs", "resolved_imbalance") g.add_edge("resolved_imbalance", "choose_best_policy") g.add_edge("choose_best_policy", "generate_explanation") g.add_edge("generate_explanation", "simulate_and_report") g.add_edge("trivial_do_nothing", "simulate_and_report") g.add_edge("simulate_and_report", END) app = g.compile(checkpointer=MemorySaver()) app here is reused for everything in the rest of this article. One graph, many independent zone stories - told apart by thread_id , not by rebuilding anything. Running Multiple Cycles MemorySaver plus a thread_id is the whole mechanism. Call .invoke() again on the same thread_id , and the graph picks up exactly where the last call left off. Only the first call needs a full initial state. Every call after that only needs whatever's new - an ops note, or nothing at all. Downtown Core, starting in deficit, no ops notes: [Cycle 1] ratio=1.93 | policy=surge_pricing | wait 6.3min โ 4.6min | resolved=NO Highest profit at $230.78, even though it did not resolve the imbalance. [Cycle 2] ratio=1.41 | policy=surge_pricing | wait 4.6min โ 4.3min | resolved=YES RESOLVED Two calls, same thread_id . The second one already knows this is cycle 2 of an ongoing deficit - nothing had to be passed back in by hand. Scheduling a Condition Across Cycles An ops note on cycle 1 says traffic will be bad for the next hour - that's 4 cycles. reconcile_inputs extracts traffic_duration_minutes=60 . Code converts that to 4 cycles, and schedules the expiry. No further notes are needed. apply_scheduled_conditions keeps traffic_level="heavy" active through cycle 4, then reverts it automatically at cycle 5 - exactly on schedule: notes = {1: "There's a bad accident on the main road backing up traffic for the next hour."} results = run_cycles(app, "downtown-traffic-duration", downtown, ops_notes=notes) cycle=1 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}} cycle=2 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}} cycle=3 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}} cycle=4 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}} cycle=5 traffic=0.00 scheduled={} One note on cycle 1 shapes five cycles of behavior. Nobody has to remind the agent about the traffic on cycles 2 through 4. A Future Trigger, and Overriding a Schedule Two more cases worth showing: A future trigger, not an immediate correction. "The concert starts at 6pm" - it's currently 5pm. That means event_flag should stay False for now, and flip to True only once that hour actually arrives. reconcile_inputs extracts event_starts_at_hour=18 . Code works out how many cycles away that is, and schedules the flip: cycle=1 hour=17 event_flag=False cycle=2 hour=17 event_flag=False cycle=3 hour=17 event_flag=False cycle=4 hour=17 event_flag=False cycle=5 hour=18 event_flag=True cycle=6 hour=18 event_flag=True A later note overriding an earlier schedule. Say a new note arrives before a scheduled condition would have expired on its own - "actually the accident's been cleared." It replaces the schedule entirely, instead of waiting for the original timer. New information wins: notes = { 1: "Traffic is bad for the next hour due to an accident.", 3: "Actually the accident has been cleared, traffic is back to normal.", } cycle=1 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}} cycle=2 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}} By cycle 2 the zone had already resolved on its own. So the schedule never actually got overridden here. But the mechanism - a new note wiping and replacing an existing schedule - is the same one that fires wh
Comments
No comments yet. Start the discussion.