Knowledge Architecture & AI Evaluation Research
Research into how structured knowledge, retrieval systems, uncertainty, governance, and human judgement shape the behaviour of AI systems.
PROTEX examines how AI systems retrieve, organise, interpret, and present knowledge when accuracy, uncertainty, source boundaries, and decision responsibility matter.
The work focuses on knowledge architecture, AI evaluation, retrieval reliability, uncertainty preservation, hallucination resistance, evidence grounding, Human-AI Collaboration, and AI governance within structured knowledge environments.
The objective is methodological: to develop research, teaching, and evaluation frameworks for understanding how knowledge architecture influences AI behaviour, rather than treating AI reliability as a model-only problem.
AI systems are increasingly used before their knowledge environments are fully understood.
Modern AI systems are connected to documents, procedures, policies, knowledge bases, SharePoint libraries, structured corpora, and operational workflows. However, many reliability problems originate not only from the model, but from the structure, quality, ownership, and evidentiary organisation of the knowledge behind it.
PROTEX asks how AI systems should be evaluated when they operate on real knowledge: fragmented documents, conflicting sources, ambiguous narratives, outdated guidance, incomplete procedures, unclear metadata, and uncertain decision boundaries.
PROTEX is supported by a working AI prototype used for research, teaching, benchmarking, and live demonstration.
Unlike purely conceptual research, PROTEX includes a functioning AI environment that can be demonstrated through a chat-based interface. The prototype allows users to observe how structured knowledge, retrieval design, uncertainty, confidence levels, and interpretation boundaries affect AI-generated answers.
The system functions as a knowledge architecture laboratory, AI evaluation environment, benchmark platform, and teaching demonstrator for explaining how modern AI systems operate on complex knowledge.
Research Environment
Used to study retrieval behaviour, evidence grounding, uncertainty handling, answer stability, and epistemic boundaries in AI-supported systems.
Teaching Demonstrator
Used in lectures and workshops to show how AI systems retrieve knowledge, respond to different types of questions, and handle ambiguity or incomplete evidence.
Evaluation Platform
Used to develop benchmark methodologies for testing factual grounding, completeness, source adherence, false-premise resistance, and uncertainty preservation.
The work examines knowledge before AI, and system behaviour after retrieval.
AI reliability depends on both the knowledge environment and the system built on top of it. The research therefore examines two complementary dimensions: whether knowledge is structured in ways that support reliable AI use, and how AI systems behave when retrieving, analysing, and reasoning over that knowledge.
AI Knowledge Readiness
This framework studies how knowledge repositories influence AI behaviour before retrieval begins.
- Knowledge source quality
- Document structure and clarity
- Metadata and findability
- Ownership and review cycles
- Conflicting or duplicated content
- Decision boundaries and escalation logic
AI Reliability Evaluation
This framework studies how AI systems behave when operating on structured or semi-structured knowledge.
- Answer accuracy and completeness
- Retrieval and source adherence
- Omission and over-interpretation risk
- False-premise resistance
- Uncertainty preservation
- Consistency across repeated questions
AI reliability failures are not limited to hallucination.
In knowledge-sensitive settings, a problematic AI answer is not always obviously false. It may be plausible, fluent, and partially grounded, while still omitting a critical condition, over-interpreting a source, blending procedures, ignoring uncertainty, or presenting unsupported confidence.
This research treats AI reliability as an interaction between model behaviour, retrieval architecture, knowledge structure, evidence quality, governance, and human responsibility.
Operational Risk
- Incorrect procedures followed
- Critical exceptions missed
- Outdated guidance reused
- Escalation steps skipped
- Incomplete answers treated as complete
Governance Risk
- Unclear accountability
- Weak evidence trails
- Insufficient human oversight
- Undefined decision boundaries
- Unsupported interpretations accepted
Knowledge Risk
- Fragmented documentation
- Conflicting sources
- Unclear source authority
- Poor metadata
- Weak uncertainty representation
What should be measured when AI operates on knowledge?
The programme investigates how AI systems behave when users ask ambiguous questions, documents conflict, procedures are incomplete, evidence is distributed across several sources, or the correct response is to refuse, qualify, or escalate.
Answer Behaviour
- Does the system answer the right question?
- Does it retrieve appropriate sources?
- Does it follow evidence rather than over-interpret it?
- Does it preserve uncertainty when sources are incomplete?
- Does it identify when it does not know?
- Does it remain consistent across repeated questions?
Knowledge Structure
- Are authoritative sources clearly identifiable?
- Are facts separated from commentary and interpretation?
- Are ownership, review cycles, and version control clear?
- Are conflicting documents detectable?
- Can omissions, exceptions, and boundaries be evaluated?
- Can answers be traced back to reliable evidence?
The research examines subtle forms of AI failure.
Visible hallucination is only one failure mode. In evidence-sensitive environments, more subtle failures may arise through omission, over-interpretation, weak grounding, uncertainty collapse, source confusion, or excessive confidence.
Omission
The system gives an answer but leaves out a condition, exception, dependency, approval step, deadline, uncertainty marker, or escalation route that changes how the answer should be understood.
Over-interpretation
The system goes beyond the available source material, turns guidance into a rule, treats commentary as procedure, or presents judgement as established fact.
Weak grounding
The system relies on sources that are outdated, secondary, incomplete, contradictory, or not authoritative for the question being answered.
Structured evaluation rather than informal impressions.
The research uses structured testing, controlled prompts, corpus analysis, retrieval observation, false-premise testing, uncertainty evaluation, and source-boundary analysis to study AI behaviour within knowledge-intensive environments.
The aim is not to declare whether an AI system is generally “good” or “bad”, but to examine how it behaves under defined epistemic conditions: missing evidence, ambiguity, conflicting sources, unclear authority, repeated prompts, and decision-boundary pressure.
Retrieval Behaviour
Examination of how systems retrieve, prioritise, omit, combine, or overextend information from available knowledge sources.
- Factual accuracy
- Completeness
- Omission behaviour
- Source adherence
- Answer consistency
Knowledge Architecture
Study of how document structure, metadata, provenance, and evidentiary organisation influence AI responses.
- Document structure
- Information architecture
- Metadata strategy
- Knowledge source quality
- Uncertainty preservation
Governance & Boundaries
Analysis of responsibility structures, decision boundaries, escalation routes, and human oversight in AI-supported knowledge systems.
- Knowledge ownership
- Decision boundaries
- Escalation logic
- Human oversight
- Responsibility mapping
Outputs from the programme are methodological, analytical, educational, and research-oriented.
The programme produces research papers, benchmark findings, methodological frameworks, evaluation criteria, case observations, teaching materials, and publications concerning AI reliability and knowledge architecture.
Knowledge Readiness Research
- Knowledge structure observations
- Source quality analysis
- Document architecture findings
- Metadata and findability observations
- Ownership and review-cycle considerations
- Conflict, duplication, and versioning research
AI Reliability Research
- Reliability evaluation criteria
- Evidence-grounding observations
- Retrieval behaviour findings
- Omission and over-interpretation analysis
- Uncertainty preservation observations
- Decision-boundary research
Studying AI systems where knowledge quality and decision boundaries matter.
The research is relevant to knowledge-intensive environments where inaccurate, incomplete, outdated, or overconfident AI answers can affect interpretation, procedure, compliance, judgement, learning, or organisational trust.
Current research contexts include Microsoft Copilot environments, Copilot Studio agents, internal knowledge assistants, procedural knowledge systems, structured behavioural case repositories, compliance knowledge bases, expert advisory systems, and operational documentation environments.
- Microsoft Copilot and Copilot Studio environments
- Internal AI assistants and agents
- Organisational knowledge repositories
- Procedural and policy documentation
- Structured behavioural and narrative knowledge corpora
- Legal, compliance, finance, and advisory knowledge systems
- Operational teams using internal procedures
- Educational demonstrations of AI system behaviour
Independent research into retrieval reliability, knowledge architecture, and evidence-bound AI.
This work is informed by PROTEX research on Microsoft Copilot Studio, retrieval reliability, corpus design, uncertainty preservation, hallucination resistance, evidence-boundary adherence, and AI behaviour within constrained knowledge environments.
The research focuses on how AI behaves when tested against structured questions, false premises, omission risks, source-boundary risks, document ambiguity, and operational knowledge constraints.
https://zenodo.org/records/20490517
https://zenodo.org/records/20431380
https://zenodo.org/records/21006312
https://protex-profiler.ai/ai-governance-framework/
View report
View report
Open to research discussion, teaching opportunities, and methodological collaboration.
PROTEX welcomes discussion with researchers, universities, students, training providers, practitioners, and domain specialists interested in AI evaluation, knowledge architecture, uncertainty preservation, AI governance, Human-AI Collaboration, and evidence-grounded AI.
Collaboration may involve research conversations, guest lectures, workshops, postgraduate teaching, methodological review, benchmark design, evaluation criteria, corpus design, governance research, working papers, or exploratory studies.
Research Contact
For research collaboration, academic discussion, teaching opportunities, methodological questions, or related enquiries: