---
schema_version: '1.0'
id: news-20260729-505933
url: https://gotosocial.chinng-lab-srv.dev/codifying-the-judge-programmatic-distillation-of-llm-evaluators/
url_hash: 505933c66a343c86eb8893e9ab6fb8d7be97fd7662a96e2d3e58751ff9c0345a
canonical_url: https://gotosocial.chinng-lab-srv.dev/codifying-the-judge-programmatic-distillation-of-llm-evaluators
source: ghost-chinng-lab
category: news/general
category_raw: LLM
region: JP
tags:
- LLM
- 評価
- 蒸留
- Programmatic Distillation
- LLM-as-Judge
- Codefy-Judge
- 機械学習
lang: ja
published_at: '2026-07-29T03:48:34Z'
fetched_at: '2026-07-29T11:35:28.765011Z'
updated_at: '2026-07-29T11:51:38Z'
status: published
content_hash: null
content_changed_at: null
license_note: full
summary: 大規模言語モデルを評価するジャッジシステムを、解釈可能で実行可能なプログラムに蒸留する手法「Programmatic Distillation」を提案。LLMジャッジの課題であるコスト・遅延・再現性の問題を解決し、教師モデルと同等以上の判断品質を維持しながら、推論コストを99%以上削減、遅延を90%以上削減することを実現。MT-Bench、RewardBench、MixEvalなど複数のベンチマークで検証され、Codefy-Judgeとして実装が公開されている。
summary_source: llm
summary_en: '"Programmatic Distillation" is a method to distill the judge system to
  an interpretable and executable program. Resolve cost, delay, and再現roducibility
  issues of LLM Judge to reduce thrust costs by more than 99% while maintaining the
  same quality as the teacher model. It has been verified by multiple command lines
  such as MT-Bench, RewardBench, MixEval, and implemented as Codefy-Judge。'
entities:
- name: LLM
  type: artifact
- name: 評価
  type: UNKNOWN
- name: 蒸留
  type: UNKNOWN
- name: Programmatic Distillation
  type: method
- name: Codefy-Judge
  type: artifact
- name: 機械学習
  type: UNKNOWN
- name: THE ANSWER
  type: organization
- name: Judge
  type: person
key_facts: []
related: []
related_auto:
- name: AIエージェント
  type: concept
  weight: 1.0
- name: Codefy_Judge
  type: content
  weight: 1.0
- name: RegressionTax
  type: concept
  weight: 1.0
- name: docker_agent
  type: organization
  weight: 1.0
- name: 生成AI
  type: artifact
  weight: 1.0
title: 'Codifying the Judge: Programmatic Distillation of LLM Evaluators'
---

# Codifying the Judge: Programmatic Distillation of LLM Evaluators

## TL;DR
大規模言語モデルを評価するジャッジシステムを、解釈可能で実行可能なプログラムに蒸留する手法「Programmatic Distillation」を提案。LLMジャッジの課題であるコスト・遅延・再現性の問題を解決し、教師モデルと同等以上の判断品質を維持しながら、推論コストを99%以上削減、遅延を90%以上削減することを実現。MT-Bench、RewardBench、MixEvalなど複数のベンチマークで検証され、Codefy-Judgeとして実装が公開されている。

## Key Points
- LLM / 評価 / 蒸留 / Programmatic Distillation / LLM-as-Judge / Codefy-Judge / 機械学習

## Details
(本文なし。リンク先参照)

## Source
元記事: [Codifying the Judge: Programmatic Distillation of LLM Evaluators](https://gotosocial.chinng-lab-srv.dev/codifying-the-judge-programmatic-distillation-of-llm-evaluators/) — published 2026-07-29T03:48:34Z
