Explore concepts like Self-Correct, Self-Refine, Self-Improve, Self-Contradict, Self-Play, and Self-Knowledge, alongside o1-like reasoning elevation🍓 and hallucination alleviation🍄.
-
Updated
Dec 7, 2024 - Jupyter Notebook
Explore concepts like Self-Correct, Self-Refine, Self-Improve, Self-Contradict, Self-Play, and Self-Knowledge, alongside o1-like reasoning elevation🍓 and hallucination alleviation🍄.
Awesome LLM Self-Consistency: a curated list of Self-consistency in Large Language Models
CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning
Zero-cost epistemic uncertainty quantification & hallucination detection for LLMs (90,000x faster than Semantic Entropy)
Black-box AI reliability certification via self-consistency sampling and conformal calibration
Capable, auditable coding that runs fully offline on a 16 GB machine. A verification-first layer (hard test execution, symbolic checking, agentic repair) that takes a local 7B to parity with its 671B teacher on verifiable tasks. MIT, pre-registered, reproducible.
The official PyTorch implementation for the Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
KG-RAG + ToT + multi-agent LLMs for evidence-grounded QA with Neo4j and fine-tuning; reproducible medical case study & evaluation.
Object oriented package for solving self-consistent mean field theory of interacting lattice systems.
GenPark AI Agent Skill - Self-consistency sampling aggregator, majority voting consensus engine, and semantic clusterer for reasoning verification.
GenPark AI Agent Skill - Self-consistency sampling aggregator, majority voting consensus engine, and semantic clusterer for reasoning verification.
Self-Consistency majority voting engine aggregating multiple stochastic reasoning paths to extract high-confidence consensus solutions.
Self-Consistency majority voting engine aggregating multiple stochastic reasoning paths to extract high-confidence consensus solutions.
Prompt-engineering study on LLM math reasoning (GSM8K) and code generation (HumanEval): zero/few-shot, self-consistency, self-verification, and experiments on prompt quality, complexity, demonstrations, and diversity.
When more sampling stops helping: a reasoning model can generate a right answer long before it can pick one. The modal and correlation ceilings of test-time scaling, with paper, figures, and code (Bay & Yearick).
Perl implementation of Markov Chain for the course BIO331
An evaluation of prompting techniques (Zero-Shot CoT, Few-Shot, Self-Consistency) on the Mistral-7B model for mathematical reasoning. This project systematically benchmarks 7 distinct methods on the GSM8K dataset.
Code and results for consistency-gated self-correction in large language model reasoning.
GSM8K-Consistency is a benchmark database for analyzing the consistency of Arithmetic Reasoning on GSM8K.
TACT: signed, label-free confidence weighting for self-consistency voting — with the thin-window boundary (2.5–7.5% of items) that explains why six other designs died
To associate your repository with the self-consistency topic, visit your repo's landing page and select "manage topics."