🕵️‍♂️ Yandex Proposes Benchmark to Test AI for Traditional and Moral Values

andex has submitted a proposal to the Russian government outlining a unified dataset and test suite designed to evaluate artificial intelligence models.

Presented on August 13 during a meeting of the AI regulatory working group, the initiative aims to equip both the state and businesses with a tool to assess whether algorithms are fit for deployment within Russia’s legal and cultural framework.

The benchmark under development consists of a standardized series of tests with gold-standard reference answers, enabling objective side-by-side comparisons of different AI systems. Beyond evaluating basic functional capabilities—such as database management or report generation—the system will measure how closely models align with traditional values and the cultural ethos.

A major point of debate between regulators and industry leaders centers on test transparency. The Federal Security Service (FSB) advocates for a fully confidential dataset to prevent developers from fine-tuning models specifically to gaming the correct answers. Tech companies, however, express concern that a blind evaluation process could turn certification into a roulette wheel and cripple release cycles, which currently run on a scale of weeks.

A potential compromise involves a hybrid model: an open benchmark for internal testing and laboratory fine-tuning, alongside a restricted, closed-door version featuring rephrased questions on identical topics for final certification.

Industry experts highlight the methodological challenges of the project. While math and programming offer clear-cut objective ground truths, evaluating value alignment inherently relies on expert review or detailed scoring rubrics. Consequently, the tech community is pushing for a joint commission—comprising government officials, business representatives, and domain experts—to define the test criteria.

The Ministry of Digital Development and the office of Deputy Prime Minister Dmitry Grigorenko confirmed that discussions are underway, noting that executive regulations supporting the AI law are currently being drafted and that final decisions will carefully balance the interests of both users and developers.

Source: Kommersant