Welcome back to Diary of an AI Architect. Today, we are stepping outside our usual solo deep dives to bring you a guest essay by my colleague Prosenjit Das, Global Black Belt for AI at Microsoft.
I’m genuinely excited to introduce Prosenjit to this community. I’ve personally learned a great deal from his expertise and guidance over the past few years at Microsoft. He works closely with Microsoft’s top enterprise customers, helping them translate Agentic AI innovation into measurable business value. He brings a unique perspective on bridging intelligent applications with real-world enterprise challenges.
My earlier posts on context engineering and decision quality and the three-layer Context Stack explored the governed data foundation agents need. Prosenjit makes that tactical in this essay, showing how Microsoft’s intelligence stack can address operational challenges.
This week in my diary, he connects Foundry IQ for document grounding, Fabric IQ for real-time and historical evidence, and Work IQ for authorized workplace context. The goal is a shared intelligence layer designed for relevant evidence, secure access and human control. Below are his words.
Table of Contents
Why build an Intelligence layer?
Who can use the Intelligence layer and how?
What are some industrial use cases and what does the Intelligence layer look like for them?
Connected Worker Safety
SportsIQ
How to build or measure one?
Below are the words of Prosenjit Das, GBB AI Apps at Microsoft | October 2026
1. Why build an Intelligence layer?
In my years of experience working as a Developer, then Enterprise Architect and currently in Microsoft and focused only on the top Global customers across several industries, I have always seen the information needed for one operational decision sits in several systems, siloed.
For example, in Construction, Manufacturing or Retail, current events arrive through cameras, sensors or transactions; history sits in analytical stores; procedures and approvals live elsewhere.
In sports, game moments must drive all decisions but key historical context lives elsewhere and company marketing strategy or policy guardrails at other places.
A human or an autonomous agent must relate these sources before it can give a useful and properly correlated answer. Repeating those joins, definitions and retrieval decisions inside each manual process or agent makes the system harder to maintain and its answers harder to explain. Also drives latency and cost considerably! Cost for siloed manual buildout or cost in terms of tokens.
Instead, imagine building an Intelligence layer to assemble context for a particular task that separates domain reasoning from retrieval orchestration. A question about one work zone in Construction or one game in Sports rarely needs the whole estate.
When a knowledge request reaches the knowledge layer, the knowledge base takes responsibility for retrieval orchestration. It can use an LLM to plan the request into focused subqueries, select the appropriate configured knowledge sources, execute searches in parallel using keyword, vector, or hybrid retrieval, semantically re-rank the results, and aggregate the most relevant evidence.
If the initial results do not meet relevance standards, it should perform iterative retrieval when configured with the appropriate reasoning effort. It then returns either the retrieved grounding content or an optionally synthesized answer, together with source citations, while enforcing the caller’s permissions.
The domain agent therefore does not need to know how individual knowledge sources should be searched. It receives the task-relevant evidence from Knowledge layer and applies domain-specific reasoning, business rules, confidence requirements, and actions to that evidence.
Source access should remain reusable and governed, so agents just call the knowledge layer using its own identity or the user’s identity and retrieve relevant information they are authorized to access rather than each rebuilding joins or getting tightly coupled with disparate knowledge sources and their drifts. The original systems still remain authoritative and control fine grained user access. This reduces avoidable context without discarding the basis for the answer.
2. Who can use the Intelligence layer and how?
Once built, the same intelligence layer can be used by different personas (and Agents) within the org.
A safety officer,
venue operator,
support specialist or
analyst
can use it through Teams or a Copilot experience. Domain agents can consume structured context through authorized APIs or Model Context Protocol (MCP) tools once they are triggered by events they have subscribed to.
They share entity definitions but receive different authorized slices and action rights. An event-triggered workflow still needs explicit delegation; the event itself cannot supply a user’s permissions.
Leveraging the intelligence layer must remain strictly decoupled from the channels where users or agents operate. But each builds their own scoped memory through each interaction with the same central knowledge layer.
Person A and Person B start with same knowledge base on day-1 but over time they build their own intelligence layer and memory scoped to each of their own.
As an example, in my IP that I built for the Sports industry, I map approved documents and policy grounding to Foundry IQ, and structured historical and real-time event data and business meaning to Fabric IQ.
Work IQ takes control in individual and collaboration context, ownership, hierarchy and prior decisions through an approved integration; Weather, traffic and schedules provide additional context and can be sourced from WebIQ or simply web search through Foundry IQ based on the exact need. The workload determines which capabilities are needed. A vector index alone cannot supply the entity relationships or action authority.
I keep the requesting person’s permissions attached to retrieval, including document access and source-level row or object restrictions. An agent identity cannot substitute for those rights. Checks on sensitive information must also follow it through retrieval, tools and the final response. The reader should be able to inspect the cited source and see where the evidence is strong and where not.
Once the Intelligence layer is built that each company can call their own IQ, they can point CoPilot or Cowork or Custom domain agents (or sub-agents) or coding agents at them to fetch the same clean and approved knowledge every single time, and over time build their own knowledge graph, hyper-personalized!
3. What are some industrial use cases and what does the Intelligence layer look like for them?
Connected Worker Safety
Business Outcome: Help safety teams spot potential hazards on a construction site and decide how to respond, while keeping people in charge of safety decisions.
In the construction-safety architecture I developed earlier this year, a planner routes real-time, historical or mixed questions. The event path runs from camera capture to object storage, a function and a vision assessment. The resulting structured event enters Fabric Real-Time Intelligence and an agent workflow. The vision output is a typed risk assessment, allowing downstream components to work with explicit observations rather than repeatedly interpreting the same image.
For a zone alert, retrieval should cover environmental and equipment events for the relevant period, then use the dimensional Lakehouse for related incidents, certifications and corrective actions. External benchmarks can add perspective without becoming the authority for the response.
Real-time and historical agents supply evidence; a synthesizer applies the current standard operating procedure (SOP), and a communication agent routes the recommendation. Authorized task and certification records may be relevant, but worker identity or certification must not be inferred from camera images.
Historical analysis adds useful context but must not delay established alerts or stop-work procedures. Those paths must remain operational if an AI dependency fails.
Human judgement and stop-work authority remain in place; the AI cannot certify safety or automatically authorize work to resume. Ontology relationships can organize the evidence, but they do not establish causality; trained predictive models and ontology-based explanations remain future work.
SportsIQ
Business Outcome: SportsIQ connects what is happening during the game with a fan’s preferences, purchase history and current venue conditions. This helps teams recommend something useful, such as an available food offer at a nearby concession during a break, while avoiding interruptions during an important play. The intended outcome is a better fan experience, more purchases that would not otherwise have happened, and fewer irrelevant or unfulfillable promotions. Teams measure success through additional profit compared with a control group, offer redemption, fulfillment, and fan opt-out rates.
For an illustrative match-day offer, I retrieve weather, ticket inventory, point-of-sale activity and authorized loyalty context for the game and seat zone in question. A versioned model of Fan, Game, Offer, Seat Zone and Sponsor makes those relationships explicit.
Eligibility rules and budget, contact-fatigue, staffing and venue-capacity limits constrain the proposal. Changing inventory needs to be rechecked before execution: an offer can be commercially sensible but operationally unsuitable if the venue cannot support the additional demand.
I separate policy interpretation from enforcement. Deterministic middleware checks the rules, and human rejection or timeout means no write, and changed inputs require reapproval. Idempotency, concurrency checks, bounded retries and a stop switch address execution failures. Durable memory also needs authorized reads and writes, retention, deletion and revocation.
4. How to build or measure one?
The starting point is one task and its accountable owner, followed by a map of entities, authoritative sources, freshness limits and access boundaries. A read-only implementation provides a baseline for retrieval and reasoning before writes are introduced.
Specify what constitutes a correct answer, an appropriate abstention and a required escalation. This makes the evaluation reflect the operational task rather than only the fluency of the response.
The measurement approach I proposed in either implementation looks somewhat like this -
Correlated traces capture model, retrieval, tool and handoff versions alongside latency and cost.
Agent Insights groups recurring issues and representative traces, with affected versions and likely causes.
A subject-matter expert should confirm or dismiss each suspected issue; the grouping helps investigation but is not a conclusive diagnosis.
Each evaluation case records whether the task was resolved, partly resolved or unresolved, whether the evidence supported the answer, and whether the right response was an answer, abstention or escalation. The reviewer identifies the failure layer, severity, confidence, and authoritative source. Reviewing a random control sample alongside selected failures helps reduce selection bias. An example enters the evaluation dataset only after explicit approval.
In this approach, the baseline is evaluated on the approved, versioned golden and regression dataset before optimization, and its manifest is frozen for reproducibility. Engineers can then change the model, prompt, knowledge, tool or agent topology and compare the candidate on the same tasks and datasets. Results need examination by case, domain and severity across quality, safety, outcome, latency and cost, because an overall improvement can hide a regression in a critical task.
A human release decision precedes shadow operation or a limited release ring, followed by monitoring and a decision to promote or roll back. Production must not automatically rewrite or promote itself. Domain experts own authoritative content, expected behavior and risk thresholds; engineers own the experiments. Managed platform tooling supports the work, while the organization retains the release decision.
A specialist agent is worth testing when traces reveal distinct tools or permissions, knowledge authority, state or memory, concentrated failures or costs, or different safety criteria. I would compare a single agent, a specialist and a stateless MCP tool under the same evidence and access requirements.
Dividing every business domain into another agent adds coordination work without necessarily improving the task. Also adds dynamic trajectory that makes evaluation harder.
For worker safety use case, I assess intent accuracy, query and data correctness, SOP alignment, relevance and actionability, using Open Telemetry traces to locate failures. Tests should include stale telemetry, noisy sensors, revoked access, missing procedures, cross-domain memory requests and attempts to bypass approval.
Rejected, expired or altered commands must not execute. Missed hazards and false positives need separate reporting because they have different operational consequences.
Context efficiency can be assessed through relevant retrieval, unnecessary prompt tokens, repeated tool calls, and latency and cost per successfully completed, authorized task.
Replaying cases with recorded source versions helps distinguish a context defect from a model change. I would accept lower token use only if the essential evidence and safe behavior are preserved, with the domain owner agreeing the trade-offs and release thresholds.
If a safety answer cites an obsolete procedure, I would investigate source selection and freshness before changing the model. If a SportsIQ offer fails because inventory changed, I would examine the execution-time check. This separation matters: the next experiment should address the component that failed, and the regression case should preserve enough evidence to show whether the change actually helped.
I am aware it’s a lot to achieve for organizations who have to responsibly manage real production systems but Foundry’s inbuilt Eval harness (Traces, Agent Insights, Evaluation and Optimizer) makes it a lot easier for organizations.
Across these implementations, my focus remains on giving people and agents relevant, current context within clear access and action boundaries. Operational evidence and domain review should guide how that context and the architecture supporting it improves over time.
About the guest
Editor’s bio: Prosenjit Das is a Global Black Belt for Azure Apps and AI at Microsoft. His work focuses on intelligent applications for enterprise customers, with interests spanning multi-agent architecture.
Connect with Prosenjit on LinkedIn.
♻️ If this was useful, share it with someone building with AI.
✉️ Subscribe at newsletter.karuparti.com so you never miss an edition.
P.S. Want more? 👋
1/ My visual guide to agentic AI → Gumroad
2/ Daily deep dives on agentic AI architecture → LinkedIn
5/ Visual frameworks and carousels → Instagram
6/ 60-second production lessons → TikTok
References
Editor-added sources support platform capabilities, not independent verification of private implementations.
Anu Karuparti Creator, Diary of an AI Architect How enterprises actually ship AI to production
Read by 3,000+ FDEs, AI Architects, and Engineering Leaders from Microsoft, Google, IBM, PwC and others.
Connect with me on LinkedIn.
Want to partner? Email me at anurag.karuparti@gmail.com.







