Compare baseline and candidate RAG outputs using lexical metrics. Get per-case score changes, regression flags, and a pass/fail threshold decision for up to 100 test cases. No LLM calls. Experimental: does not verify factual accuracy or semantic correctness.
Diagnose Shopify feed disagreements with original HTTP offers and rendered variant, currency, price and stock evidence. Captures subscription purchase context, screenshots and load failures. Experimental beta with explicit theme selectors; unknown evidence stays inconclusive.
Check saved dataset timestamps before RAG ingestion. Evaluate up to 1,000 metadata records; get batch pass/fail and per-record diagnostics. No crawling or truth verification. n8n tutorial: https://recency-n8n-guide.saimislam.chatgpt.site