sakutto
生成AI

Amazonが希少本を裁断してAI学習データにしている

Amazon学習データ著作権
Amazonが希少本を裁断してAI学習データにしている

追跡調査が希少本の到着先を突き止めた

きっかけはAI企業がどこから学習データを仕入れているのかという疑問でした。取材の方法がそのまま答えになっています。

希少本に追跡装置を仕込んだ

404 Media が追跡装置を仕込んだのはAI企業に買われそうだと踏んだ希少本でした。その荷物が国内を移動していく様子を追いました。Amazon の書籍買い付けの実態が報じられたのはこれが初めてだと 404 Media は書いています。 学習データの調達は契約や社内の話に閉じがちで、外から観測できる形になった例は多くありません。

公式情報を見る →
"A 404 Media investigation was able to reveal Amazon's book buying operation, which hasn't been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination."— 404 Media の調査報道より

行き先はラスベガスの施設VGT3だった

追跡装置が最後にたどり着いたのはネバダ州ラスベガスにある Amazon の倉庫でした。この場所で働くチームは VGT3 と呼ばれています。シンボルは歯をむき出しにして本を手にした恐竜です。 何をする場所なのかはマークが正直に語っています。

公式情報を見る →
"That final destination was an Amazon warehouse in Las Vegas, Nevada."(到着先)/"The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands."(VGT3 チーム)— 404 Media の調査報道より

製本を外して裁断しスキャンする

倉庫で何が行われているのか。ここが今回いちばん生々しい部分です。

従業員が作業の中身を語った

ここで働く従業員は届いた大量の印刷本の製本を切り落とし、速くスキャンできる状態にするのが仕事だと話しています。理由として語られているのは速くスキャンするためで、紙の本はその過程で失われます。 保管でも再販でもなく、読み取ったら終わりです。買い集めているのが希少本である以上、失われるのは替えの利く在庫ではありません。

公式情報を見る →
"Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process."— 404 Media の調査報道より

Amazonはどこまで説明したか

Amazon は 404 Media に対し、顧客が使う製品とサービスを改善するために商業ルートで書籍を購入していると述べています。買っていることは認めていますが、裁断やスキャンについては触れていません。 用途は「製品とサービスの改善」だけ。どのモデルの学習に使うのか、集めた本文をどう保管するのかは説明されていません。

公式情報を見る →
"The facility, known as VGT3, identifies itself with a symbol of a dinosaur holding a book in its claws. Amazon told 404 Media in a statement that it \"purchases books through commercial channels to improve the products and services customers use.\""— TechCrunch より

なぜ紙の希少本が狙われるのか

読み終えて残るのは「なぜ本なのか」という疑問です。理由は学習データ側の事情にあります。

2022年より前の文章が学習データとして効く

大規模言語モデルの学習には途方もない量の文章が要ります。インターネットから取れるものはすでに取り尽くされました。そこで浮上したのが絶版になっていたりネットでは見つからなかったりする希少本です。2022年より前に出版された文章には「AIが書いた可能性がない」という確実さがあります。 生成AIが普及したあとのテキストにこの保証は付きません。

公式情報を見る →
"Companies like Amazon need unfathomably large amounts of text to train their LLMs, which have already ingested what they can from the internet (and, in Anthropic's case, illegally pirated books). Rare books, especially ones that are out of print or impossible to find on the internet, offer a new source of coveted training data."— TechCrunch より

モデル崩壊を避けるための素材になる

AIが生成した文章を大量に取り込むと、出力の質が落ちていく「モデル崩壊」が起こりえます。学習の材料が自分たちの出力で汚染されていく構図です。人の手で書かれたと確実に言える素材はそれを避けるための資源になります。紙にしか残っていない文章が学習データとして重宝されるのは、そのためです。

公式情報を見る →
"These texts are especially valuable since there's no chance that anything published before 2022 was written by an LLM. When LLMs train on AI-generated text, they risk \"model collapse,\" which can occur when the quality of an LLM's outputs degrade after ingesting too much AI-generated text."— TechCrunch より

著作権と学習データで何が争点か

学習データの調達をめぐる争いはすでに各所で起きています。TechCrunch は記事の中で、Anthropic が違法に海賊版の書籍を取り込んだ事例に触れています。正規に買った本を裁断してスキャンする行為は、少なくとも入手の経路としては海賊版とは別物です。 ただし買ったからといって学習に使ってよいことになるのかは、権利者の側から見れば別の問いです。事業者に何をどこまで開示させるかという枠組みはEU AI法の8月施行カリフォルニアのAI透明化法でも動いている最中です。

公式情報を見る →
"Companies like Amazon need unfathomably large amounts of text to train their LLMs, which have already ingested what they can from the internet (and, in Anthropic's case, illegally pirated books)."— TechCrunch より

調査報道の元記事は途中から有料会員限定になりますが、本記事の事実関係は無料で読める範囲だけで確かめられます。ページを構造ごとマークダウンへ落としておけば、見出しと箇条書きを保ったまま引用の突き合わせに使えます。

無料ツールURLマークダウン変換URL(ウェブページ)を入力するだけでマークダウン(Markdown)に変換。見出し・表・リスト・リンクを保持したままmd化でき、LLMやRAGの前処理、調査資料の整形にも最適な無料オンラインツール。今すぐ使ってみる →

まとめ

希少本1冊に追跡装置を入れる。それだけの方法で、AIの学習データが調達される現場がひとつ表に出ました。行き先はラスベガスの VGT3、作業は製本を切り落としてのスキャン、紙の本はその過程で失われます。Amazon は購入の事実だけを認め、裁断には触れていません。背景にあるのは「AIが書いていないと確実に言える文章」の価値の高さで、紙にしか残っていない文章ほど狙われる構図になっています。 学習データの出どころをどこまで開示させるのかという議論は、これから制度の側で詰まっていくところです。

よくある質問

Q. この件はどうやって明らかになったのですか?
404 MediaがAI企業に買われそうな希少本へ追跡装置を仕込み、国内を移動する荷物を追いかけました。到着先はラスベガスにあるAmazonの倉庫です。Amazonの書籍買い付けの実態が報じられたのはこれが初めてだと404 Mediaは書いています。
404 Media — We Tracked a Shipment of Rare Books
A 404 Media investigation was able to reveal Amazon's book buying operation, which hasn't been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination. 404 Media — We Tracked a Shipment of Rare Books
Q. 本はスキャンした後どうなるのですか?
壊れます。従業員の話では速くスキャンするために製本を切り落としており、その過程で紙の本は失われます。返品や再販を前提とした扱いではありません。
404 Media — We Tracked a Shipment of Rare Books
Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. 404 Media — We Tracked a Shipment of Rare Books
Q. なぜ古い紙の本が学習データとして価値を持つのですか?
2022年より前に出版された文章はAIが書いたものではないと確実に言えるからです。AIが生成した文章を学習し続けると出力の質が落ちる「モデル崩壊」が起こりえます。だからAI企業は人が書いたと分かる素材を探しています。
TechCrunch — Amazon, which started off selling books, is destroying rare texts to train AI
These texts are especially valuable since there's no chance that anything published before 2022 was written by an LLM. When LLMs train on AI-generated text, they risk "model collapse," which can occur when the quality of an LLM's outputs degrade after ingesting too much AI-generated text. TechCrunch — Amazon, which started off selling books, is destroying rare texts to train AI

関連ツール

関連ツールカテゴリ

記事