Engineering LeadershipHands-On AI QAQuality EngineeringTest Automation & Release AssuranceConnect on LinkedIn

PARIMI · KNOWLEDGE BASE

Practical knowledge for engineering reliable AI systems.

Deep technical guides, implementation patterns and engineering perspectives on AI quality, agent testing, LLM evaluation, RAG, Cognigy and Evaluation-Driven Development.

Search the library

Find practical guides by problem, technology, testing method or AI quality topic.

Knowledge library

Guides, articles and engineering notes

7 articles
AI QA Engineering

Sep 24, 2026

How to Validate an LLM-as-a-Judge for AI Agent Evaluation

A research-backed engineering method for validating an LLM judge before using it to evaluate AI agent answers, tool calls, trajectories and release decisions.

#LLM-as-a-Judge#Agent-Evaluation#AI-QA#Evals
AI QA Engineering

Sep 23, 2026

AI Agent Regression Testing with Playwright: From Browser Automation to Evidence-Driven QA

A production architecture for combining Playwright execution, browser traces, deterministic assertions and AI evaluation without turning Playwright into an opaque LLM judge.

#Playwright#AI-QA#Regression#Automation
AI Agent QA

Sep 23, 2026

How to Test an AI Agent: A Research-Backed Evaluation and Assurance Framework

A rigorous engineering framework for evaluating agentic systems across task success, trajectories, tools, state, semantic quality, safety, statistical reliability and regression.

#AI-QA#AI-Agents#Evaluation#Agentic-Systems
AI Security Testing

Sep 23, 2026

Prompt Injection Testing for AI Agents: Threat Modeling, Exploitation and Control Validation

A security-engineering treatment of direct and indirect prompt injection, excessive agency, data exfiltration and tool-boundary testing for agentic systems.

#AI-Security#Prompt-Injection#Agent-Security#Red-Teaming
RAG & LLM Evaluation

Sep 23, 2026

RAG Evaluation at Research Depth: Retrieval, Grounding, Faithfulness and Hallucination

A rigorous framework for diagnosing RAG systems across retrieval quality, context utility, claim-level faithfulness, completeness, abstention and production drift.

#RAG#LLM-Evaluation#Grounding#Hallucination
Cognigy

Sep 23, 2026

Testing Cognigy Agents: From Flow Semantics to Trace-Level Assurance

A research-informed testing architecture for Cognigy agents covering flows, intents, nodes, state, tool/API behaviour, handoffs, multilingual cases and regression evidence.

#Cognigy#Conversational-AI#Agentic-AI#AI-QA
Agentic AI Testing

Sep 23, 2026

Testing Tool Calling and Agent Orchestration: Contracts, Trajectories and Failure Semantics

A deep engineering guide to evaluating tool choice, argument correctness, sequencing, authorization, retries, orchestration and side effects in agentic systems.

#Agentic-AI#Tool-Calling#Orchestration#AI-QA