Cross-provider AI code review for Claude Code — evidence-based confidence scoring with Codex, Gemini & Claude
-
Updated
Jul 19, 2026 - Shell
Cross-provider AI code review for Claude Code — evidence-based confidence scoring with Codex, Gemini & Claude
Uncertainty based selection of compatible inputs
Extract structured data from any document — PDF, DOCX, HTML, CSV, plain text — using LLMs with Pydantic schema validation, per-field confidence scores, and source grounding.
Open-source LLM evaluation engine with statistical confidence scoring
Zero-Noise utilities for safer product research and review signal analysis.
Multi-agent AI task delegation architecture for n8n: orchestrator routes natural-language commands to specialist agents with confidence scoring and human-in-the-loop gates.
System that aggregates outputs from multiple Large Language Models (GPT-4, Claude-3, custom models) to generate reliable, high-confidence results through consensus-based reasoning evaluation. Demonstrates sophisticated AI orchestration with 92.7% accuracy improvement over single-model.
Research-grade Self-Correcting RAG agent built with LangGraph that retrieves knowledge, generates answers, evaluates grounding/relevance/completeness, and iteratively self-improves with confidence scoring and memory.
Deterministic structured extraction from noisy LLM/OCR output. Zero LLM round-trips, microsecond latency, confidence score on every result. msgspec · Pydantic · dataclasses.
Verification system that catches coding agents falsely claiming task completion. Runs 4 parallel checks (file integrity, test quality, scope narrowing, optional LLM judge) over task+claim+diff and returns a weighted 0-100 confidence score with evidence.
AI-powered concierge that normalises guest messages from WhatsApp, Booking.com, Airbnb, Instagram and direct channels, drafts a reply with Claude, and routes responses through a deterministic confidence-scoring pipeline. Built with FastAPI + Claude Sonnet 4.
Runtime reliability intelligence designed specifically for OpenClaw frameworks and agents.
Smart Document Conversion for the AI Era - CPU-only, fast, with confidence scoring. Converts PDF, DOCX, PPTX, HTML, EPUB to Markdown, JSON, HTML, Text.
Audit template for AI video-understanding timestamp confidence, clip evidence, speaker intent, source context, and edit decisions.
7-axis weighted confidence function for AI output quality. Evidence, reasoning, calibration, source, domain, coherence, meta.
Enterprise-grade Confidence-Driven State Reconciliation Platform leveraging Kafka Streams, Redis, PostgreSQL, and Spring Boot to preserve competing truths, compute evidence-backed consensus, deterministic replay, and operational analytics.
Multi-source candidate data transformer — merges Resume, CSV, ATS, LinkedIn, GitHub & Notes into a conflict-resolved canonical profile with confidence scoring and provenance tracking.
Scorecards for AI video-understanding clip context, timestamp confidence, speaker intent, edit rationale, and source evidence.
RAG service with a policy engine that decides whether to answer, ask a clarifying question, or refuse before generation ever runs. Lexical retrieval, SQLite-backed chunk storage, page-level citations, and a mocked generation layer shaped for a drop-in real LLM call.
Self-healing and drift recovery for agents. A zero-runtime-dependency TypeScript library and CLI that scores output confidence, detects ungrounded claims with GSAR-style typed grounding, monitors behavioral drift, and orchestrates recovery (rollback, retry, escalate, ask a human), with a tamper-evident audit trail.
Add a description, image, and links to the confidence-scoring topic page so that developers can more easily learn about it.
To associate your repository with the confidence-scoring topic, visit your repo's landing page and select "manage topics."