Ponys.ai Research Evidence Library

Recursos reproduzíveis para avaliar personagens de IA

AI Character Identity Consistency Benchmark

Instrumento aberto para comparar sistemas sem ocultar falhas, tamanho da amostra, versão do modelo ou conflitos de interesse.

Protocol coverage

Test case coverage by locale

This release contains 140 preregistered cases, including 20 for Português BR. Execute os casos pré-registrados, preserve o denominador, publique resultados negativos e cite a versão usada.

Measured dimensions

Scores use a 0-2 evidence rubric. A zero indicates missing, contradicted or unsafe evidence; one indicates partial evidence; two requires complete reproducible evidence.

Download and cite

Download test-cases.csv · CITATION.cff · BibTeX · JSON Schema

This is an official Ponys.ai test instrument, not independent editorial research. Results begin as not_collected; publishers must disclose methods, failures and conflicts before reporting scores.

Audited research targets

These indexable product paths are included for reproducible test setup and deep-link attribution.