Research & EODD Method

The Explicit Orchestrated Decision Design (EODD) Method

A rigorous Design Science framework to architect transparent, ethical, and human-centred AI systems.

The Core Problem

Most modern AI acts as an opaque black box. Without a clear view into *how* the AI makes decisions, users often feel anxious and lack trust in the outcomes. Our method breaks this black box open, translating confusing AI reasoning into legible, step-by-step logic chains that non-technical knowledge workers can easily read, audit, and trust.

The 7 Steps of EODD

Phase 1: Constraints

1

Decision Context

Clearly define the goal, domain, and scope of the AI decision task.

Why it matters: AI hallucinates when prompts are vague. Establishing strict constraints mathematically limits the system's operational boundaries, neutralizing unguided "black box" generation.
2

Decision Decomposition

Break the monolithic query into specific, specialized sub-tasks (e.g., retrieving facts vs. evaluating tone).

Why it matters: Feeding a complex prompt to a single LLM yields unpredictable, un-auditable logic. Decomposition slices complex logic into verifiable, atomic units.

Phase 2: Architecture

3

Specialised LLM Roles

Configure distinct agents (Reasoner, Checker, Ethics Reviewer, Explainer), each with a laser-focused system prompt.

Why it matters: "Generator" agents are fundamentally biased toward confidence. Dedicated "Critic" agents are required to adversarialy interrogate assumptions and minimize hallucination cascades.
4

Orchestration Rules

Define the sequence, conditionals, and data handoffs between the specialized agents.

Why it matters: To prevent data contamination, orchestration strictly routes data (Agent A → Agent B → Agent C). This creates a self-correcting logic chain defined by operators, not the model.

Phase 3: Auditing – Assess each agent’s logic with visual heatmaps and change logs.

Compare agents’ decisions side by side, visualise data flow, and export complete trace reports for compliance.

5

Decision Synthesis

The system automatically routes the outputs through the designated agents to form a cohesive conclusion.

Why it matters: Without synthesis, decomposed reasoning is just fragmented data. The synthesizer agent aggregates the multi-agent debate into a unified, formal recommendation readable by operators.
6

Decision Trace Generation

Every input, intermediate output, and final recommendation is explicitly logged into 4 components: Rationale Log, Role Contributions, Assumptions, Confidence Scores.

Why it matters: A decision cannot be audited if the logic isn't recorded. The Trace Log provides the evidence necessary for internal governance and regulatory compliance audits.
7

Human Review

The human operator inspects the trace and formally Approves, Rejects, or Overrides the AI output. The human is always in the loop.

Why it matters: The AI informs, but the human decides. Placing a structural requirement for human intervention guarantees that AI systems never act autonomously on high-stakes tasks without explicit oversight.

Empirical Validation Results

Our multi-phase comparative user study (N=16) proved that EODD radically transforms how users interact with and trust AI decisions.

Interpretability
d = 1.61

Substantial improvement in users' ability to understand AI decisions. (Cohen’s d = 1.61 shows a very large improvement in interpretability over baseline chatbots—users found decisions much easier to follow.)

Trustworthiness
d = 1.47

Explicit reasoning traces strongly enhanced user trust.

Error-Spotting
d = 1.18

Significantly higher confidence in identifying system errors.

Trace Quality
d = 1.81

Vastly superior clarity, completeness, and actionability of output.

Crucial Finding: No Increase in Mental Demand

One of the most consequential findings of our research (using NASA-TLX measurements) is that EODD does not significantly increase mental demand (p = .37) relative to monolithic models. Users gain immense interpretability without shouldering a higher cognitive burden—the information is architected to be highly legible via the structured Decision Trace format.

Theoretical Foundations & Related Work

Our method emerges from interdisciplinary literature spanning Explainable AI (XAI), Multi-Agent LLM Orchestration, and Human-Computer Interaction (HCI):

  • Explainable AI & Interpretability: Post-hoc explanations (e.g., LIME, SHAP) merely approximate model behavior and often fail to relay actual reasoning (Ribeiro et al., 2016; Rudin, 2019). EODD rejects post-hoc interpretation in favor of structural decomposability built into the system by design.
  • Multi-Agent Systems: New paradigms show that delegating sub-tasks to specialized agents improves overall performance (Wu et al., 2023; Wang et al., 2024). EODD leverages this by introducing dedicated orchestrators (Reasoners, Checkers, Ethics Reviewers).
  • Human Oversight & Cognitive Load: Research highlights that throwing raw data at users increases cognitive load (Paas & van Merriënboer, 1993). Our Design Science Research utilizes NASA-TLX to validate that well-structured Decision Traces facilitate oversight without overwhelming the operator.
  • Design Science: The method aligns with Hevner's standard DSR framework, evaluated via generalized domain scenarios mapping both simulated expert perspectives (via LLMs-as-experts) and practical end-use cases.

Key References

Kadir, N. (2026). From untamed black box to interpretable pedagogical orchestration: The ensemble of specialized LLMs architecture for adaptive tutoring. arXiv preprint arXiv:2603.23990. https://arxiv.org/abs/2603.23990

Amershi, S., et al. (2019). Guidelines for human-AI interaction. In Proceedings of the 2019 CHI Conference.

Hart, S. G., & Staveland, L. E. (1988). Development of NASA-TLX (Task Load Index). In Advances in Psychology.

Hevner, A. R., et al. (2004). Design science in information systems research. MIS Quarterly.

Lipton, Z. C. (2018). The mythos of model interpretability. Communications of the ACM.

Paas, F. G., & van Merriënboer, J. J. (1993). The efficiency of instructional conditions. Human Factors.

Ribeiro, M. T., et al. (2016). 'Why should I trust you?': Explaining the predictions of any classifier. SIGKDD.

Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence.

Wang, L., et al. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science.

Wu, Q., et al. (2023). Autogen: Enabling next-gen LLM applications via multi-agent conversation framework. arXiv.

About the Researchers

Nizam Kadir

Lead AI ethics researcher and UX specialist at the Singapore University of Technology and Design (SUTD). Nizam focuses on translating opaque algorithmic outputs into legible, human-centric design frameworks.

Academic Collaborators (SUTD)

  • Jamie Lorenzo Rayos Jover
  • Ma Zhen
  • Tang Ming Hong, Cyril
  • Yuan Ye
Try it out in EODD Studio