PARIMI · KNOWLEDGE BASE
Practical knowledge for engineering reliable AI systems.
Deep technical guides, implementation patterns and engineering perspectives on AI quality, agent testing, LLM evaluation, RAG, Cognigy and Evaluation-Driven Development.
Search the library
Find practical guides by problem, technology, testing method or AI quality topic.
Knowledge library
Guides, articles and engineering notes
Sep 24, 2026
How to Validate an LLM-as-a-Judge for AI Agent Evaluation
A research-backed engineering method for validating an LLM judge before using it to evaluate AI agent answers, tool calls, trajectories and release decisions.
Sep 23, 2026
AI Agent Regression Testing with Playwright: From Browser Automation to Evidence-Driven QA
A production architecture for combining Playwright execution, browser traces, deterministic assertions and AI evaluation without turning Playwright into an opaque LLM judge.
Sep 23, 2026
How to Test an AI Agent: A Research-Backed Evaluation and Assurance Framework
A rigorous engineering framework for evaluating agentic systems across task success, trajectories, tools, state, semantic quality, safety, statistical reliability and regression.
Sep 23, 2026
Prompt Injection Testing for AI Agents: Threat Modeling, Exploitation and Control Validation
A security-engineering treatment of direct and indirect prompt injection, excessive agency, data exfiltration and tool-boundary testing for agentic systems.
Sep 23, 2026
RAG Evaluation at Research Depth: Retrieval, Grounding, Faithfulness and Hallucination
A rigorous framework for diagnosing RAG systems across retrieval quality, context utility, claim-level faithfulness, completeness, abstention and production drift.
Sep 23, 2026
Testing Cognigy Agents: From Flow Semantics to Trace-Level Assurance
A research-informed testing architecture for Cognigy agents covering flows, intents, nodes, state, tool/API behaviour, handoffs, multilingual cases and regression evidence.
Sep 23, 2026
Testing Tool Calling and Agent Orchestration: Contracts, Trajectories and Failure Semantics
A deep engineering guide to evaluating tool choice, argument correctness, sequencing, authorization, retries, orchestration and side effects in agentic systems.