---
schema_version: '1.0'
id: news-20260719-e1a06a
url: https://gotosocial.chinng-lab-srv.dev/in-place-tokenizer-expansion-for-pre-trained-llms-duo-yan-yu-dekodogao-su-hua-noshi-jian-de-kuo-zhang/
url_hash: e1a06af33104042f924cdab6513f70854b52571ebe7d2f483a6fd1049abed421
canonical_url: https://gotosocial.chinng-lab-srv.dev/in-place-tokenizer-expansion-for-pre-trained-llms-duo-yan-yu-dekodogao-su-hua-noshi-jian-de-kuo-zhang
source: ghost-chinng-lab
category: news/tech
category_raw: AI・テクノロジー
region: JP
tags:
- AI・テクノロジー
- 言語モデル
- 自動拡張
- 訓練データ
- LLMs
- メモリ使用
lang: ja
published_at: '2026-07-19T09:41:04Z'
fetched_at: '2026-07-19T12:38:30.693005Z'
updated_at: '2026-07-19T12:38:47Z'
status: published
content_hash: null
content_changed_at: null
license_note: full
summary: 事前学習済みLLMのトークナイザーを、既存トークンを保持しつつ新規トークンを追加する拡張手法。新規トークンは元のサブトークン埋め込みの平均で初期化、二段階（Embedding訓練→全体継続学習）で適用。ヒンディー語などで2.2〜4.0倍のデコード速度改善を実現、多言語対応のコスト削減を可能にする。
summary_source: llm
summary_en: Extended method to add new s while holding existing s. The new  is initialized
  at the average of the original sub-token embedded, and applied in two phases (Embedding
  training → overall co ity learning). Achieve 2.2-4.0 sampling speed improvement
  in Hindi, etc.,  cost reduction for multilingual support。
entities:
- name: 大規模言語モデル
  type: artifact
- name: LLMs
  type: concept
- name: computational_resource_expansion
  type: method
- name: FORBES JAPAN
  type: organization
key_facts: []
related: []
related_auto:
- name: Nvidia
  type: organization
  weight: 1.0
- name: agent_crew_3_bot
  type: person
  weight: 1.0
- name: 計算資源
  type: data
  weight: 1.0
- name: AutoMem
  type: concept
  weight: 1.0
- name: 記憶管理
  type: concept
  weight: 1.0
title: In-Place Tokenizer Expansion for Pre-trained LLMs — 多言語デコード高速化の実践的拡張
---

# In-Place Tokenizer Expansion for Pre-trained LLMs — 多言語デコード高速化の実践的拡張

## TL;DR
事前学習済みLLMのトークナイザーを、既存トークンを保持しつつ新規トークンを追加する拡張手法。新規トークンは元のサブトークン埋め込みの平均で初期化、二段階（Embedding訓練→全体継続学習）で適用。ヒンディー語などで2.2〜4.0倍のデコード速度改善を実現、多言語対応のコスト削減を可能にする。

## Key Points
- AI・テクノロジー / 言語モデル / 自動拡張 / 訓練データ / LLMs / メモリ使用

## Details
(本文なし。リンク先参照)

## Source
元記事: [In-Place Tokenizer Expansion for Pre-trained LLMs — 多言語デコード高速化の実践的拡張](https://gotosocial.chinng-lab-srv.dev/in-place-tokenizer-expansion-for-pre-trained-llms-duo-yan-yu-dekodogao-su-hua-noshi-jian-de-kuo-zhang/) — published 2026-07-19T09:41:04Z
