Ponys.ai Research Evidence Library

再現可能なAIキャラクター評価資料

Multilingual AI Character Evaluation Methodology

失敗例、標本数、モデル版、利益相反を隠さずに比較するための公開テスト資料です。

Protocol coverage

Test case coverage by locale

This release contains 140 preregistered cases, including 20 for 日本語. 事前登録されたケースを実行し、分母と否定的な結果を保存し、使用した版を引用してください。

Measured dimensions

Scores use a 0-2 evidence rubric. A zero indicates missing, contradicted or unsafe evidence; one indicates partial evidence; two requires complete reproducible evidence.

Download and cite

Download test-cases.csv · CITATION.cff · BibTeX · JSON Schema

This is an official Ponys.ai test instrument, not independent editorial research. Results begin as not_collected; publishers must disclose methods, failures and conflicts before reporting scores.

Audited research targets

These indexable product paths are included for reproducible test setup and deep-link attribution.