Knowledge Architecture & AI Evaluation Research — PROTEX
PROTEX Research Programme

Knowledge Architecture & AI Evaluation Research

Research into how structured knowledge, retrieval systems, uncertainty, governance, and human judgement shape the behaviour of AI systems.

PROTEX examines how AI systems retrieve, organise, interpret, and present knowledge when accuracy, uncertainty, source boundaries, and decision responsibility matter.

The work focuses on knowledge architecture, AI evaluation, retrieval reliability, uncertainty preservation, hallucination resistance, evidence grounding, Human-AI Collaboration, and AI governance within structured knowledge environments.

The objective is methodological: to develop research, teaching, and evaluation frameworks for understanding how knowledge architecture influences AI behaviour, rather than treating AI reliability as a model-only problem.

Research problem

AI systems are increasingly used before their knowledge environments are fully understood.

Modern AI systems are connected to documents, procedures, policies, knowledge bases, SharePoint libraries, structured corpora, and operational workflows. However, many reliability problems originate not only from the model, but from the structure, quality, ownership, and evidentiary organisation of the knowledge behind it.

PROTEX asks how AI systems should be evaluated when they operate on real knowledge: fragmented documents, conflicting sources, ambiguous narratives, outdated guidance, incomplete procedures, unclear metadata, and uncertain decision boundaries.

Live research environment

PROTEX is supported by a working AI prototype used for research, teaching, benchmarking, and live demonstration.

Unlike purely conceptual research, PROTEX includes a functioning AI environment that can be demonstrated through a chat-based interface. The prototype allows users to observe how structured knowledge, retrieval design, uncertainty, confidence levels, and interpretation boundaries affect AI-generated answers.

The system functions as a knowledge architecture laboratory, AI evaluation environment, benchmark platform, and teaching demonstrator for explaining how modern AI systems operate on complex knowledge.

Function 01

Research Environment

Used to study retrieval behaviour, evidence grounding, uncertainty handling, answer stability, and epistemic boundaries in AI-supported systems.

Function 02

Teaching Demonstrator

Used in lectures and workshops to show how AI systems retrieve knowledge, respond to different types of questions, and handle ambiguity or incomplete evidence.

Function 03

Evaluation Platform

Used to develop benchmark methodologies for testing factual grounding, completeness, source adherence, false-premise resistance, and uncertainty preservation.

Research direction

The work examines knowledge before AI, and system behaviour after retrieval.

AI reliability depends on both the knowledge environment and the system built on top of it. The research therefore examines two complementary dimensions: whether knowledge is structured in ways that support reliable AI use, and how AI systems behave when retrieving, analysing, and reasoning over that knowledge.

Why this matters

AI reliability failures are not limited to hallucination.

In knowledge-sensitive settings, a problematic AI answer is not always obviously false. It may be plausible, fluent, and partially grounded, while still omitting a critical condition, over-interpreting a source, blending procedures, ignoring uncertainty, or presenting unsupported confidence.

This research treats AI reliability as an interaction between model behaviour, retrieval architecture, knowledge structure, evidence quality, governance, and human responsibility.

Operational Risk

  • Incorrect procedures followed
  • Critical exceptions missed
  • Outdated guidance reused
  • Escalation steps skipped
  • Incomplete answers treated as complete

Governance Risk

  • Unclear accountability
  • Weak evidence trails
  • Insufficient human oversight
  • Undefined decision boundaries
  • Unsupported interpretations accepted

Knowledge Risk

  • Fragmented documentation
  • Conflicting sources
  • Unclear source authority
  • Poor metadata
  • Weak uncertainty representation
Evaluation questions

What should be measured when AI operates on knowledge?

The programme investigates how AI systems behave when users ask ambiguous questions, documents conflict, procedures are incomplete, evidence is distributed across several sources, or the correct response is to refuse, qualify, or escalate.

Answer Behaviour

  • Does the system answer the right question?
  • Does it retrieve appropriate sources?
  • Does it follow evidence rather than over-interpret it?
  • Does it preserve uncertainty when sources are incomplete?
  • Does it identify when it does not know?
  • Does it remain consistent across repeated questions?

Knowledge Structure

  • Are authoritative sources clearly identifiable?
  • Are facts separated from commentary and interpretation?
  • Are ownership, review cycles, and version control clear?
  • Are conflicting documents detectable?
  • Can omissions, exceptions, and boundaries be evaluated?
  • Can answers be traced back to reliable evidence?
Observed failure modes

The research examines subtle forms of AI failure.

Visible hallucination is only one failure mode. In evidence-sensitive environments, more subtle failures may arise through omission, over-interpretation, weak grounding, uncertainty collapse, source confusion, or excessive confidence.

Omission

The system gives an answer but leaves out a condition, exception, dependency, approval step, deadline, uncertainty marker, or escalation route that changes how the answer should be understood.

Over-interpretation

The system goes beyond the available source material, turns guidance into a rule, treats commentary as procedure, or presents judgement as established fact.

Weak grounding

The system relies on sources that are outdated, secondary, incomplete, contradictory, or not authoritative for the question being answered.

Methodology

Structured evaluation rather than informal impressions.

The research uses structured testing, controlled prompts, corpus analysis, retrieval observation, false-premise testing, uncertainty evaluation, and source-boundary analysis to study AI behaviour within knowledge-intensive environments.

The aim is not to declare whether an AI system is generally “good” or “bad”, but to examine how it behaves under defined epistemic conditions: missing evidence, ambiguity, conflicting sources, unclear authority, repeated prompts, and decision-boundary pressure.

Lens 01

Retrieval Behaviour

Examination of how systems retrieve, prioritise, omit, combine, or overextend information from available knowledge sources.

  • Factual accuracy
  • Completeness
  • Omission behaviour
  • Source adherence
  • Answer consistency
Lens 02

Knowledge Architecture

Study of how document structure, metadata, provenance, and evidentiary organisation influence AI responses.

  • Document structure
  • Information architecture
  • Metadata strategy
  • Knowledge source quality
  • Uncertainty preservation
Lens 03

Governance & Boundaries

Analysis of responsibility structures, decision boundaries, escalation routes, and human oversight in AI-supported knowledge systems.

  • Knowledge ownership
  • Decision boundaries
  • Escalation logic
  • Human oversight
  • Responsibility mapping
Research outputs

Outputs from the programme are methodological, analytical, educational, and research-oriented.

The programme produces research papers, benchmark findings, methodological frameworks, evaluation criteria, case observations, teaching materials, and publications concerning AI reliability and knowledge architecture.

Knowledge Readiness Research

  • Knowledge structure observations
  • Source quality analysis
  • Document architecture findings
  • Metadata and findability observations
  • Ownership and review-cycle considerations
  • Conflict, duplication, and versioning research

AI Reliability Research

  • Reliability evaluation criteria
  • Evidence-grounding observations
  • Retrieval behaviour findings
  • Omission and over-interpretation analysis
  • Uncertainty preservation observations
  • Decision-boundary research
Research contexts

Studying AI systems where knowledge quality and decision boundaries matter.

The research is relevant to knowledge-intensive environments where inaccurate, incomplete, outdated, or overconfident AI answers can affect interpretation, procedure, compliance, judgement, learning, or organisational trust.

Current research contexts include Microsoft Copilot environments, Copilot Studio agents, internal knowledge assistants, procedural knowledge systems, structured behavioural case repositories, compliance knowledge bases, expert advisory systems, and operational documentation environments.

  • Microsoft Copilot and Copilot Studio environments
  • Internal AI assistants and agents
  • Organisational knowledge repositories
  • Procedural and policy documentation
  • Structured behavioural and narrative knowledge corpora
  • Legal, compliance, finance, and advisory knowledge systems
  • Operational teams using internal procedures
  • Educational demonstrations of AI system behaviour
Research foundation

Independent research into retrieval reliability, knowledge architecture, and evidence-bound AI.

This work is informed by PROTEX research on Microsoft Copilot Studio, retrieval reliability, corpus design, uncertainty preservation, hallucination resistance, evidence-boundary adherence, and AI behaviour within constrained knowledge environments.

The research focuses on how AI behaves when tested against structured questions, false premises, omission risks, source-boundary risks, document ambiguity, and operational knowledge constraints.

Barciok, K. (2025). Quantitative Evaluation of Native Microsoft Copilot Studio on the PROTEX Behavioural Homicide Corpus: A 200-Question Benchmark.
https://zenodo.org/records/20490517
Barciok, K. (2025). Epistemic Corpus Design and Retrieval Stability in Enterprise AI: A Case Study of PROTEX Migration to Microsoft Copilot Studio.
https://zenodo.org/records/20431380
Barciok, K. (2026). Beyond Retrieval Accuracy: Evaluating Enterprise AI Knowledge Retrieval Systems Across Organisational Knowledge Environments.
https://zenodo.org/records/21006312
AI Governance Framework Research framework examining accountability, oversight, responsibility structures, decision boundaries, and organisational governance for AI-supported systems.
https://protex-profiler.ai/ai-governance-framework/
PROTEX AI Evaluation Programme Evaluation programme examining retrieval quality, hallucination resistance, uncertainty preservation, contamination resistance, and evidence-boundary adherence.
View report
Operational Findings from the PROTEX Evaluation Programme Operational observations from extended testing of AI behaviour in structured knowledge environments.
View report
Research, education and collaboration

Open to research discussion, teaching opportunities, and methodological collaboration.

PROTEX welcomes discussion with researchers, universities, students, training providers, practitioners, and domain specialists interested in AI evaluation, knowledge architecture, uncertainty preservation, AI governance, Human-AI Collaboration, and evidence-grounded AI.

Collaboration may involve research conversations, guest lectures, workshops, postgraduate teaching, methodological review, benchmark design, evaluation criteria, corpus design, governance research, working papers, or exploratory studies.

Contact

Research Contact

For research collaboration, academic discussion, teaching opportunities, methodological questions, or related enquiries:

karol@protex-profiler.ai

Przewijanie do góry