All Articles

Permission-Aware Context Retrieval: How Unblocked, Glean, Augment Code, and Sourcegraph Compare

Four AI context platforms compared on query-time permission enforcement, deployment options, and audit logs, with a 7-point checklist for the vendor security review.

Permission-Aware Context Retrieval: How Unblocked, Glean, Augment Code, and Sourcegraph Compare

The short version: permission-aware context retrieval comes in two shapes. Query-time enforcement checks the asker's current entitlement when the question is asked; index-time enforcement checks a copy captured at crawl time. Unblocked's Data Shield checks per query across code and non-code sources, Glean mirrors ACLs into its index and filters against the synced copy, Sourcegraph mirrors code-host permissions for Cody with polling and optional webhooks, and Augment Code documents repository-level scoping at install rather than per-user permission sync. Audit logs range from Enterprise-plan CSV exports (Unblocked) to JSON on stderr (Sourcegraph) to a feature still marked in progress (Augment Code).

Permission-aware context retrieval is the property that decides whether an AI coding agent can leak code or documents its user was never allowed to open. The four platforms most often shortlisted for it (Unblocked, Glean, Augment Code, and Sourcegraph) enforce permissions in different places, at different times, across different sources. Those differences matter more than any feature list. A Cloud Security Alliance and Token Security survey of 418 IT and security professionals found that 65% of organizations had an AI-agent-related incident in the prior 12 months. In 61% of those cases the incident involved data exposure (Cloud Security Alliance, April 2026). This comparison covers where each platform enforces permissions, how it syncs them, where it can be deployed, and what its audit logs record, so an RBAC- or SOC 2-driven review can be run against facts rather than brochures.

What is query-time vs. index-time permission enforcement?#

Query-time permission enforcement checks the requesting user's current access in the source system (or a continuously mirrored copy of it) at the moment a question is asked. Any candidate document the user cannot open is dropped before the model sees it. Index-time enforcement records each document's permissions when the content is crawled, stores that snapshot alongside the index, and filters results against the snapshot. They diverge the moment access changes: a revoked permission that lingers in a snapshot is a live leak until the next crawl, while a query-time check catches it on the next question.

The distinction matters because the model cannot be trusted to filter. AWS's security team put it plainly in August 2026: the agent "acts as an orchestrator, not a gatekeeper," and "authorization is enforced by downstream services," so that "even if the agent is compromised through prompt injection or application bugs, it can't access unauthorized data" (AWS Security Blog, August 2026). A March 2026 arXiv study measured the gap: under a permissive policy, social engineering succeeded against the model 74.6% of the time; under a deterministic pre-action policy, attackers achieved a 0% success rate (arXiv 2603.20953, March 2026).

How do Unblocked, Glean, Augment Code, and Sourcegraph compare?#

The four platforms below are not interchangeable. Two are context engines for coding agents, one is enterprise search, and one is code search with an AI assistant attached. Each cell reflects the vendor's own documentation as of September 2026; "not documented" means the docs do not state it, which is itself a finding for a security review.

PlatformPrimary Use CasePermission SyncDeployment OptionsAudit Logs
UnblockedContext engine for engineering teams and coding agents across code, PRs, Slack, Jira, Confluence, Notion, and 20+ sourcesData Shield checks the asker's access in the source system per query; identity inherited through SAML SSO, SCIM, and OAuth-linked accounts; per-source toggle, off by default except SlackCloud; on-premises or own-cloud on Enterprise; SOC 2 Type 2, CASA Tier 2, GDPREnterprise plan; timestamp, actor, and change description per event across data-source, CI, team-setting, and question events; searchable, filterable, CSV export; admin-only under RBAC
GleanEnterprise search and assistant across workplace apps, with MCP access for agentsConnectors sync each source's permissions map and identity data into the Knowledge Graph; results filtered against the synced copy; incremental plus periodic full identity crawls; MCP runs as the signed-in userGlean-hosted SaaS or customer-hosted on GCP or AWSAdmin audit logs cover configuration changes only, exclude end-user activity, 30-day default retention, CSV export; separate MCP activity logs by server, tool, user, and date
Augment CodeAI coding agent with a Context Engine over code, available in IDE, CLI, cloud, and via MCPRemote Context Engine indexes selected repos' default branches chosen at GitHub App install; per-user permission sync across sources not documented; tool permissions govern agent actions rather than dataAugment-hosted cloud; self-hosted daemons run agents on your infrastructure while LLM calls still go to Augment; SOC 2 Type 2 and ISO/IEC 42001 per its security pageMarked in progress on the admin checklist; hooks and tool-permission settings let teams log decisions themselves
Sourcegraph (Cody)Code search and AI assistant across repositoriesMirrors repository permissions from GitHub, GitLab, Bitbucket, Perforce, Gerrit, and Azure DevOps; other hosts via explicit permissions API; polling sync with webhooks recommended to cut lag; requires an acls-enabled license; site admins bypass checks by defaultSourcegraph Cloud (managed), self-hosted Kubernetes via Helm, or Docker ComposeStructured JSON to stderr covering authentication, authorization, repository access, and GraphQL; several categories off by default; retention not documented

Best for: Unblocked, teams whose agents need code and non-code context under one per-query permission check. Glean, companies standardizing enterprise search across workplace apps with a synced-ACL model. Augment Code, teams that want an agent with a code index and accept repo-level scoping at install. Sourcegraph, organizations that need code-host-faithful permissions on search and Cody and are comfortable self-hosting.

What is Unblocked's approach to permission-aware context retrieval?#

Unblocked treats permission-aware context retrieval as a per-query check rather than a property of the index. Its Data Shield feature "restricts data access in Unblocked according to the permissions established in your connected third-party systems," so answers "only contain source material that the user posing the question is authorized to access" (Unblocked docs). Identity comes from the systems you already run: SAML SSO with Okta, Google Workspace, Microsoft Entra ID, AWS Identity Center, and Ping, plus SCIM provisioning, and per-user OAuth links for tools like GitHub and Jira (Unblocked docs). If an identity cannot be matched, the next question asks the user to connect that account. Coverage is explicit: 20+ sources support Data Shield, while Datadog, Sentry, Snowflake, external websites, and Stack Overflow do not, and the toggle is off by default for every source except Slack. A reviewer can see exactly which sources still need the checkbox flipped.

Unblocked is SOC 2 Type 2, CASA Tier 2, and GDPR compliant, stores no passwords, runs no identity service of its own, and never trains on customer data (Unblocked security). Enterprise-plan audit logs record the timestamp, actor, and a description of each change across data-source, CI, team-setting, and question events, and export to CSV (Unblocked docs). Deployment is cloud by default, with on-premises or own-cloud options on Enterprise (Unblocked AI info). Clio's application security lead described the practical effect:

The biggest thing was democratizing access to our GitHub repo without giving every one of our go-to-market users a GitHub Enterprise account. It wasn't even the cost, it was the management and the risk. Giving a safe way to access the codebase was the biggest thing I didn't see any other tool support.

Alec Robins — CTO & Co-founder, Rally

For how the same per-query check keeps agents from mixing sources with different rules, see centralizing context for coding agents and the Data Shield announcement.

Does Glean enforce permissions the same way?#

Not quite. Glean's version runs on mirrored ACLs. Per its documentation, connectors "sync access control information from sources (for example, ACLs or equivalent models) so retrieval respects source visibility," and "results are filtered to what the current user is allowed to see in the source, according to synced permissions" (Glean docs). That is a sync-then-filter model: the query-time check runs against Glean's copy, and freshness depends on crawl cadence, which Glean documents as incremental identity crawls plus periodic full crawls. It claims near-real-time permission updates for some connectors; ask which of your sources qualify and what the lag is for the rest.

Glean's MCP path is well documented. Every search, chat, and document call "enforces per-document and per-object permissions based on Glean's Knowledge Graph," and the MCP server "executes all tools as that specific user." It "cannot elevate privileges beyond the underlying Glean user identity" (Glean docs). Identity arrives via OIDC or SAML. Audit logs are narrower than the permission story. Admin audit logs "do not include end-user activity," retain 30 days by default, and export to CSV. MCP activity logs are tracked separately by server, tool, user, and date (Glean docs). Deployment is Glean-hosted SaaS or customer-hosted on GCP or AWS.

Where does Augment Code stand on permission sync and audit logging?#

Augment Code's documentation describes permissions for what the agent may do, and says much less about what it may read. Its remote Context Engine, reachable over MCP, indexes "selected repos' default branches (chosen during GitHub App install)," with authentication via OAuth or API key (Augment docs). Scoping therefore happens at install, per repository, and the docs do not describe per-user permission sync across sources. Tool permissions are granular and can be committed to a repo to bind every cloud agent, but they gate shells, file edits, and MCP servers rather than data access (Augment docs).

On audit logging, the enterprise admin checklist marks the feature as under construction, alongside SCIM, which Augment says it "will soon rollout" (Augment docs). Its security page lists SOC 2 Type 2 as attested and ISO/IEC 42001 as certified. Self-hosted daemons run agents on your machines, but agents "still make all LLM calls to Augment's external services," which matters when a reviewer asks where code travels (Augment docs).

How does Sourcegraph handle repository permissions for Cody?#

Sourcegraph has the most mature code-host permission model of the four and the narrowest scope. Per its documentation, repository permissions are mirrored from GitHub, GitLab, Bitbucket, Perforce, Gerrit, and Azure DevOps; any other host is supported only through the explicit permissions API. The instance "needs to be configured with a license that has acls feature enabled" (Sourcegraph docs). Cody inherits that model: "strict permissions are enforced to ensure that only code that the user has read permission for is retrieved" (Sourcegraph docs). Admins can also set Context Filters to exclude repositories from ever reaching a third-party LLM.

The freshness caveats are documented. Sync runs user-centric and repo-centric polling, and worst-case lag equals a full sync cycle. Sourcegraph's own example puts that at roughly 25 days for a large deployment, which is why it recommends "configuring webhooks for permissions on GitHub" (Sourcegraph docs). Site admins bypass permission checks by default unless authz.enforceForSiteAdmins is set. Audit logs are "structured logs delivered as JSON to STDERR" covering authentication, authorization, repository access, and GraphQL, with several categories off by default and no retention policy documented (Sourcegraph docs). Deployment is Sourcegraph Cloud, Kubernetes via Helm, or Docker Compose. Everything outside the repository, from Slack threads to Jira tickets, sits outside this model.

Which platform fits a SOC 2 or RBAC-driven security review?#

A context engine has six requirements to meet: unified context across code and non-code sources, conflict resolution, targeted retrieval, data governance, token optimization, and personalized relevance. Run the four platforms through that lens and data governance has to hold across every source the agent reads. Otherwise the agent leaks through the source with the weakest rule. Sourcegraph governs code faithfully and stops at the repository boundary. Glean governs workplace documents through synced ACLs and treats code as one connector among many. Augment Code governs what the agent does more precisely than what it reads. Unblocked applies one per-query check across code and the non-code systems where design decisions live, which is the shape a coding-agent review usually needs; conflicting context tools explains why the multi-source case is where inconsistent rules bite.

Gartner expects the average Fortune 500 enterprise to run more than 150,000 agents by 2028, up from fewer than 15 in 2025 (Gartner, April 2026). It also predicts that over half of successful attacks on AI agents will exploit access-control weaknesses and prompt injection by 2029 (Gartner, August 2026). NIST's February 2026 concept paper on agent identity and authorization asks for exactly the evidence a SOC 2 auditor will: identification, authorization, and "auditing and non-repudiation" for agent actions (NIST NCCoE, February 2026). For a wider field than these four, see the engineering knowledge platform comparison.

Frequently asked questions#

Is permission-aware context retrieval the same as RBAC? No. RBAC assigns permissions to roles and is usually checked when a user or app authenticates. Permission-aware retrieval applies to what an agent pulls into context at query time, scoped to the acting user. It can consume RBAC data from source systems, but it enforces it at a different, later point: the moment of retrieval. Authorization vendors describe the same gap from their side; WorkOS notes that per-agent roles lead to "role explosion" and that agents "need resource-level permissions, not just tenant-wide roles" (WorkOS, February 2026).

Why is retrieval-time filtering better than ingestion-time permissioning? Ingestion-time permissioning freezes access at index time, so it drifts out of date whenever someone gains or loses access. Retrieval-time filtering checks live permissions at the instant of the query, so revocations take effect immediately and the agent never assembles an answer from documents the current user cannot open.

Does permission-aware retrieval help with SOC 2? Yes. It maps onto SOC 2's CC6 access-control expectations by enforcing least privilege on what agents read, and, when paired with per-retrieval logging, it supplies the timestamped, attributable audit evidence auditors request. It is a meaningful control contribution, though full compliance spans many controls. Confirm scope with your auditor.

Can it use our existing identity provider? Yes. A well-built permission-aware system integrates with your existing identity provider and source-system groups rather than replacing them. It reads entitlements from those systems and enforces them at retrieval, so you keep one source of truth for identity while adding a filtering checkpoint agents did not previously have.

What is query-time permission enforcement? Query-time permission enforcement checks the requesting user's current access in the source system at the moment a question is asked, and removes any candidate document they cannot open before the model reads it. It differs from index-time enforcement, which checks a permissions snapshot captured when the content was crawled and can lag behind revocations.

Does Unblocked support SOC 2 Type 2? Yes. Unblocked's security page lists AICPA SOC 2 Type 2, CASA Tier 2, and GDPR compliance, encryption in transit and at rest with customer-specific keys, and no training on customer data. Authentication is delegated to your identity provider through SAML SSO or OAuth 2.0; Unblocked stores no passwords.

How does Unblocked handle audit logging? Audit logs are available on the Enterprise plan under Settings, Team Settings, Audit Logs. Each event records a timestamp, the actor, and a description of the change, across data-source, CI, team-setting, and question events. Logs are searchable, filterable by date, actor, and category, and exportable as CSV. With RBAC enabled, only admins can view them.

What should you check in a vendor security review?#

Take these seven checks into the call, in this order, and write down the vendor's answer to each rather than accepting a certification badge as a proxy. They track the ingredients AWS's prescriptive guidance for agentic AI lists for access control: "identity propagation from users through agent chains, permission boundaries for agent actions and tool access, audit trails of agent decisions and actions" (AWS Prescriptive Guidance).

  1. Enforcement point: is the permission check applied per query against the source system, or against a snapshot taken at crawl time?
  2. Identity propagation: does the agent carry the asking user's identity from SSO or OAuth into every retrieval, or does one integration credential serve everyone?
  3. Revocation lag: when access is removed in GitHub or Confluence, how long until the agent stops seeing that content, and is the number documented?
  4. Source coverage: which connected sources are outside the permission model entirely? Every vendor here has at least one.
  5. Audit granularity and retention: do logs capture who asked and what changed, are end-user events included, how long are they kept, and can you export them? A July 2026 arXiv gateway design keeps a hash-chained audit log the agent cannot alter, a useful bar to hold vendor logs against (arXiv 2607.05518, July 2026).
  6. Data path: where do prompts and code travel for inference, including in "self-hosted" modes?
  7. Evidence: ask for the SOC 2 Type 2 report itself, then map its scope to the five answers above.

A platform that answers all seven in writing is one your auditor can work with, and the answers are what permission-aware context retrieval means in practice. For the fuller definition behind that bar, start with what a context engine is. Unblocked was built to pass this review across every system your engineers already use, with the permission check attached to the question rather than to the index.