Most agent demos assume one user, one context, one source of truth. Drop that same agent into an organization and the interesting problem stops being reasoning — it’s that knowledge is split across people who don’t all have the right to see each other’s. Org-Agent treats that split as the main event instead of an afterthought.
The framing I like: before an action can proceed, the agent has to satisfy constraints it didn’t generate — who’s asking, what they’re allowed to know, whose information is still valid, and how to reconcile two users who want contradictory things. That’s not a prompt-engineering problem. It’s an access-control problem wearing an LLM.
- 🎯 Constraints come first: identity, authority, and permissions gate what evidence an action is even allowed to touch
- 🔍 Cross-user memory is the risk surface — one user’s context bleeding into another’s answer is a breach, not a hallucination
- 💡 Tasks become a dependency graph, scheduled by topological sort, so the agent gathers only what it’s permitted to before it acts
- ⚡ Resolved at query time, not by pre-indexing all of history — 44% accuracy at 3.9x fewer tokens than baselines
- 📊 Two benchmarks (MUSES-Bench, GroupMemBench) separate the “coordinate across users” skill from the “remember across users” skill
The token efficiency is the quiet headline. Constraint-centric scheduling means Org-Agent never builds a giant memory graph up front — it resolves each request through a per-task graph, which is exactly the tradeoff you want when one agent fields thousands of org requests a day. The HF paper page carries the ablations if you want to see how much the dependency modeling actually buys.
Here’s what I keep circling back to: every enterprise agent eventually hits this wall, and most teams reach for RBAC bolted onto the tool layer. Is a constraint graph inside the agent’s reasoning loop a better place to enforce who-can-know-what — or just a harder place to audit it?