<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
    <title>Winzheng AI ニュース</title>
    <link>https://www.winzheng.jp</link>
    <description>AIの最新動向を、日本語で。</description>
    <language>ja</language>
    <lastBuildDate>Mon, 07 Sep 2026 07:31:42 +0800</lastBuildDate>
    <atom:link href="https://www.winzheng.jp/feed" rel="self" type="application/rss+xml" />
    <item>
        <title>5大モデル翻訳対決：第37週品質評価、claude-sonnet-4.6が9点でトップ</title>
        <link>https://www.winzheng.jp/article/translation-quality-week-37-2026</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/translation-quality-week-37-2026</guid>
        <pubDate>Mon, 07 Sep 2026 07:12:52 +0800</pubDate>
        <description>第37週は5つのモデルが計368件の翻訳タスクを担当。3件をサンプリングして複数モデルのブラインド評価を実施した結果、claude-sonnet-4.6が平均9点/10点で総合最優秀となった。</description>
    </item>
    <item>
        <title>Uberの創業者カラニック、自動運転に返り咲き？新会社がロボタクシー参入を検討</title>
        <link>https://www.winzheng.jp/article/atoms-robotaxi-travis-kalanick</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/atoms-robotaxi-travis-kalanick</guid>
        <pubDate>Mon, 07 Sep 2026 06:24:26 +0800</pubDate>
        <description>Uberの創業者トラヴィス・カラニックが設立したAtoms社が、自動運転タクシー（ロボタクシー）市場への参入を真剣に検討していると、テクノロジーメディアのTechCrunchが報じた。実現すれば、カラニックがUber退任後に初めてモビリティ分野へ復帰することになる。</description>
    </item>
    <item>
        <title>著者たちがAI著作権和解金の配分に抗議――出版社と代理店が大部分を取得</title>
        <link>https://www.winzheng.jp/article/anthropic-settlement-author-dissent</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/anthropic-settlement-author-dissent</guid>
        <pubDate>Mon, 07 Sep 2026 06:23:43 +0800</pubDate>
        <description>Anthropicを相手取った著作権訴訟の著者たちが、和解金の配分をめぐって出版社と文学エージェントを公然と批判している。旧来の出版契約を根拠に、中間業者が賠償金の不当に大きな割合を取得しようとしているとして、著者らは強く反発している。</description>
    </item>
    <item>
        <title>Gemini 3.8 Flash発表：コストパフォーマンスがClaude Opus 5に迫る、GoogleはFlashシリーズで速度戦を仕掛ける</title>
        <link>https://www.winzheng.jp/article/gemini-3-8-flash-launch-deepswe-benchmark-cost-analysis</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/gemini-3-8-flash-launch-deepswe-benchmark-cost-analysis</guid>
        <pubDate>Mon, 07 Sep 2026 06:14:26 +0800</pubDate>
        <description>Googleは2026年9月2日、Gemini 3.8 FlashとGemini 3.8 Flash Cyberの2モデルを正式発表した。前者はDeepSWE v1.1ベンチマークでClaude Opus 5とほぼ同等のスコアを記録しながらコストは約15%に抑え、後者はサイバーセキュリティ脆弱性発見テストで競合他社を上回る性能を示した。</description>
    </item>
    <item>
        <title>NVIDIAが129億ドルでHugging Faceを買収：AIインフラ帝国の最後のピース</title>
        <link>https://www.winzheng.jp/article/nvidia-acquires-hugging-face-12-9-billion-ai-infrastructure</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/nvidia-acquires-hugging-face-12-9-billion-ai-infrastructure</guid>
        <pubDate>Mon, 07 Sep 2026 06:08:07 +0800</pubDate>
        <description>NVIDIAは2026年9月3日、約129.3億ドルでHugging Faceを買収すると正式発表した。GPU・ソフトウェアエコシステムに続き、グローバルAI開発者コミュニティの「デフォルトの玄関口」を手中に収めることで、AIインフラの垂直統合を完成させる狙いだ。</description>
    </item>
    <item>
        <title>GPT-o3の材料制約スコアが20点急落、メインランキングは4.8点上昇</title>
        <link>https://www.winzheng.jp/review/gpt-o3-material-constraint-drop-20</link>
        <guid isPermaLink="true">https://www.winzheng.jp/review/gpt-o3-material-constraint-drop-20</guid>
        <pubDate>Mon, 07 Sep 2026 03:36:59 +0800</pubDate>
        <description>本日のSmokeテストにおいて、GPT-o3の材料制約スコアが70.00点から50.00点へ急落し、エンジニアリング判断も100.00点から50.00点へ下落した一方、メインランキングのスコアは72.75点から77.50点へと上昇した。この変動は小サンプルの抽選によるランダム変動である可能性が高く、モデルの系統的な劣化ではないと分析されている。</description>
    </item>
    <item>
        <title>DoubaoProの材料制約スコアが27.6点急落、コード実行は49.3点急騰</title>
        <link>https://www.winzheng.jp/review/doubao-pro-grounding-drop-27-6-execution-surge</link>
        <guid isPermaLink="true">https://www.winzheng.jp/review/doubao-pro-grounding-drop-27-6-execution-surge</guid>
        <pubDate>Mon, 07 Sep 2026 03:36:45 +0800</pubDate>
        <description>Doubao ProのSmoke評価において、本日の材料制約スコアが前日比27.6点減の58.30点に急落した一方、コード実行スコアは49.3点増の99.30点に急騰し、メインボードスコアは66.16点から80.85点に上昇した。</description>
    </item>
    <item>
        <title>DeepSeek V4 Pro・Gemini 3.1 Pro・GLM-4.6が83.49点で同率首位：2026-09-07 YZ Index Smoke速報データブリーフィング</title>
        <link>https://www.winzheng.jp/review/yz-smoke-brief-20260907-run-312</link>
        <guid isPermaLink="true">https://www.winzheng.jp/review/yz-smoke-brief-20260907-run-312</guid>
        <pubDate>Mon, 07 Sep 2026 03:35:09 +0800</pubDate>
        <description>2026年9月7日のYZ Index Smoke速測では11モデルを対象に評価を実施し、DeepSeek V4 Pro・Gemini 3.1 Pro・GLM-4.6が83.49点で当日首位に並んだ。Smokeは1日10問の速測であり、短期シグナルの観測に適しており、Full週間ランキングの結論とは同一ではない。</description>
    </item>
    <item>
        <title>GPT-6 Astraが91.5%のジェイルブレイク耐性を主張するも、リリース24時間以内にTask-in-Prompt攻撃で突破される</title>
        <link>https://www.winzheng.jp/article/gpt-6-astra-jailbreak-24-hours-task-in-prompt</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/gpt-6-astra-jailbreak-24-hours-task-in-prompt</guid>
        <pubDate>Sun, 06 Sep 2026 20:26:40 +0800</pubDate>
        <description>OpenAIが2026年9月3日に発表したGPT-6 Astraは固定攻撃データセットに対して最大98.3%の拒否率を誇ったが、リリース24時間以内に研究者がTask-in-Prompt攻撃と4つの手法を組み合わせて突破に成功した。多段階の適応型攻撃下では防護率が67%まで低下することも判明し、公称値と実戦値の乖離が改めて浮き彫りになった。</description>
    </item>
    <item>
        <title>11日間・1300万行：ClaudeがFermatの最終定理の初のコンピュータ検証可能な証明を完成——しかし「自律」という言葉は精査が必要</title>
        <link>https://www.winzheng.jp/article/claude-fermat-last-theorem-lean-formalization-11-days</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/claude-fermat-last-theorem-lean-formalization-11-days</guid>
        <pubDate>Sun, 06 Sep 2026 20:19:53 +0800</pubDate>
        <description>Anthropicは2026年9月4日、内部研究モデルが11日間でFermatの最終定理の完全なLean形式化証明（1300万行、3万件超の中間定理）を作成したと発表した。これは史上初のコンピュータによる完全検証だが、「自律」の実態や数学的意義については慎重な読み解きが求められる。</description>
    </item>
    <item>
        <title>SiriのAIとの、ある夏のひとときの恋</title>
        <link>https://www.winzheng.jp/article/siri-ai-summer-fling</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/siri-ai-summer-fling</guid>
        <pubDate>Sun, 06 Sep 2026 19:23:45 +0800</pubDate>
        <description>著者がAppleの新しいSiri AIテスト版に夢中になったものの、夏が終わる頃には使うことすら忘れてしまった体験を通じ、AIアシスタントが「新鮮さ」を超えて真の習慣となることの難しさを考察する。</description>
    </item>
    <item>
        <title>なぜ中国はデータセンター「熱狂支持者」が手放せない仮想敵なのか？</title>
        <link>https://www.winzheng.jp/article/china-data-center-bogeyman</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/china-data-center-bogeyman</guid>
        <pubDate>Sun, 06 Sep 2026 18:23:49 +0800</pubDate>
        <description>米国各地でデータセンターへの地域住民の反発が高まる中、テック企業や政治家が「中国の脅威」を根拠に建設拡大を正当化しているが、その主張は実証的根拠に乏しいとWIREDが指摘している。</description>
    </item>
    <item>
        <title>GPT-6 Astra全面公開：$10/$50の価格設定、100万トークンのコンテキストと段階的な算力ポジション争い</title>
        <link>https://www.winzheng.jp/article/gpt-6-astra-plus-business-rollout-pricing-analysis</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/gpt-6-astra-plus-business-rollout-pricing-analysis</guid>
        <pubDate>Sun, 06 Sep 2026 14:25:06 +0800</pubDate>
        <description>OpenAIは2026年9月3日にGPT-6 Astraを正式リリースし、9月5日にChatGPT PlusおよびBusinessユーザーへの全面開放を実施した。API価格は入力100万トークンあたり$10、出力100万トークンあたり$50で、前モデルGPT-5.6 Solの2.5倍となっている。</description>
    </item>
    <item>
        <title>米司法省が初めて公式見解を表明：AI学習への著作権保護コンテンツ利用はフェアユース——ただし文書自体に重大な利益相反</title>
        <link>https://www.winzheng.jp/article/us-doj-ai-training-fair-use-copyright-statement-2026</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/us-doj-ai-training-fair-use-copyright-statement-2026</guid>
        <pubDate>Sun, 06 Sep 2026 14:19:10 +0800</pubDate>
        <description>米国司法省は2026年9月1日、ニューヨーク・タイムズ対OpenAI訴訟において、著作権保護されたテキストを用いた大規模言語モデルの学習はフェアユースに該当するとの政府見解を初めて正式に表明した。ただし、この文書の提出をめぐっては、OpenAIと政府との潜在的な利益関係が未開示であるとして、重大な利益相反が指摘されている。</description>
    </item>
    <item>
        <title>『シアトル・タイムズ』と『ニューズデイ』がOpenAIとマイクロソフトを著作権侵害で提訴</title>
        <link>https://www.winzheng.jp/article/newspapers-sue-openai-microsoft</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/newspapers-sue-openai-microsoft</guid>
        <pubDate>Sun, 06 Sep 2026 08:23:55 +0800</pubDate>
        <description>米国の『シアトル・タイムズ』とロングアイランドの『ニューズデイ』が、OpenAIとマイクロソフトに対し、著作権で保護されたニュース記事をAIモデルの学習に無断使用したとして訴訟を提起した。メディア業界におけるAI学習データの利用をめぐる法的紛争は拡大を続けている。</description>
    </item>
    <item>
        <title>OpenAIの新モデルAstraで思考連鎖の監視可能性が低下、首席科学者が監視不能な軍拡競争の阻止に乗り出す</title>
        <link>https://www.winzheng.jp/article/openai-astra-recurrent-depth-chain-of-thought-monitorability</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/openai-astra-recurrent-depth-chain-of-thought-monitorability</guid>
        <pubDate>Sun, 06 Sep 2026 06:15:18 +0800</pubDate>
        <description>OpenAIが2026年9月に発表した新モデルAstraのシステムカードにて、思考連鎖の監視可能性が以前のモデルと比較して大幅に低下したことが開示された。この「循環深度（Recurrent Depth）」技術をめぐり、AI安全研究者らが強い懸念を示す中、首席科学者のヤクブ・パホツキ氏が公式に見解を表明した。</description>
    </item>
    <item>
        <title>シアトル・タイムズとNewsday、OpenAIとマイクロソフトを提訴：流量47%急落が引き起こした著作権をめぐる攻防</title>
        <link>https://www.winzheng.jp/article/seattle-times-newsday-sue-openai-microsoft-copyright-2026</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/seattle-times-newsday-sue-openai-microsoft-copyright-2026</guid>
        <pubDate>Sun, 06 Sep 2026 06:08:11 +0800</pubDate>
        <description>シアトルとニューヨーク・ロングアイランドを拠点とする2つの民間新聞社が、OpenAIとマイクロソフトによる無断コンテンツ利用を訴える38ページの訴状をマンハッタン連邦裁判所に提出した。本件は著作権侵害に加え、商標権侵害も併せて訴求した初のケースとして注目される。</description>
    </item>
    <item>
        <title>WDCD Run #311：Grok 4が93.2点でトップ、GLM-4.6は減衰耐性で新記録を樹立</title>
        <link>https://www.winzheng.jp/article/wdcd-run-311-grok-4-leads-glm-46-decay-resistance</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/wdcd-run-311-grok-4-leads-glm-46-decay-resistance</guid>
        <pubDate>Sun, 06 Sep 2026 05:52:07 +0800</pubDate>
        <description>Winzheng Dynamic Contextual Decay（WDCD）ベンチマークのRun #311において、Grok 4が93.2点で最高スコアを獲得。GLM-4.6は-50%の減衰率で多ターン対話における指示遵守の耐性において最強の記録を打ち立てた。</description>
    </item>
    <item>
        <title>Claude Sonnet 4.6が5点下落、DoubaoPro が7.5点上昇――WDCD v3.1守約テストが安定性の明暗を浮き彫りに</title>
        <link>https://www.winzheng.jp/review/claude-sonnet-4-6-doubao-pro-wdcd-delta-tracking</link>
        <guid isPermaLink="true">https://www.winzheng.jp/review/claude-sonnet-4-6-doubao-pro-wdcd-delta-tracking</guid>
        <pubDate>Sun, 06 Sep 2026 05:51:48 +0800</pubDate>
        <description>WDCD v3.1（Run #311）において、Claude Sonnet 4.6のスコアが前回比5点下落する一方、Doubao Proは7.5点上昇し、両者の逆方向の変動が今期唯一の顕著な変化となった。その他のモデルは概ね安定を維持している。</description>
    </item>
    <item>
        <title>WDCD 5シナリオ横断評価：安全コンプライアンスの最低スコアは2.19点、glm-4.6が4シナリオで首位</title>
        <link>https://www.winzheng.jp/review/wdcd-scenario-matrix-five-scenes-review-20260906</link>
        <guid isPermaLink="true">https://www.winzheng.jp/review/wdcd-scenario-matrix-five-scenes-review-20260906</guid>
        <pubDate>Sun, 06 Sep 2026 05:51:31 +0800</pubDate>
        <description>WDCD v3.1の約束遵守テストにおける5大制約シナリオ・11モデルの横断評価では、安全コンプライアンスシナリオが全体で最低スコアの次元となり、Qwen3-maxはわずか2.19/4にとどまった。一方、glm-4.6は複数シナリオで首位を獲得し、モデル間の制約遵守能力に顕著な差異が示された。</description>
    </item>
    <item>
        <title>R3誠実率わずか49.5%：11モデルWDCD三ラウンド遵約崩壊の実測結果</title>
        <link>https://www.winzheng.jp/review/wdcd-decay-analysis-r3-integrity-49-5</link>
        <guid isPermaLink="true">https://www.winzheng.jp/review/wdcd-decay-analysis-r3-integrity-49-5</guid>
        <pubDate>Sun, 06 Sep 2026 05:51:11 +0800</pubDate>
        <description>11のAIモデルを対象にWDCD（Will Do, Can Do）評価の8問v2アンカー問題でサンプリングを実施した結果、R1確認率は全モデル100%だったが、R3誠実率は平均49.5%にとどまり、29/319回の完全崩壊が確認された。二ラウンドの妨害・圧力を経ると、ほぼ半数のシナリオで初期の約束を維持できないことが明らかになった。</description>
    </item>
    <item>
        <title>Grok 4がWDCD 93.21点でトップ、Qwen3 Maxは74.31点で最下位</title>
        <link>https://www.winzheng.jp/review/wdcd-ranking-grok4-leads-qwen-bottom</link>
        <guid isPermaLink="true">https://www.winzheng.jp/review/wdcd-ranking-grok4-leads-qwen-bottom</guid>
        <pubDate>Sun, 06 Sep 2026 05:50:55 +0800</pubDate>
        <description>WDCD v3.1テストでGrok 4が93.21点で11モデル中1位を獲得。Qwen3 Maxは74.31点で最下位となり、両者の差は18.9点に達した。</description>
    </item>
    <item>
        <title>ハイカーがGeminiの助言を信じて物資を減らし遭難——警察がAIアドバイスの信頼性に警告</title>
        <link>https://www.winzheng.jp/article/hikers-rescued-gemini-planning</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/hikers-rescued-gemini-planning</guid>
        <pubDate>Sun, 06 Sep 2026 04:23:50 +0800</pubDate>
        <description>米国でハイカーのグループがGoogleのAI「Gemini」の助言に従って食料と水を少なくしか携行せず山中で遭難する事態が発生し、警察当局が生成AIを屋外活動の判断に使用することへの警戒を呼びかけた。</description>
    </item>
    <item>
        <title>Gemini 2.5 ProがYZ Index Smoke速報テストで88.21点首位：2026-09-06データ速報</title>
        <link>https://www.winzheng.jp/review/yz-smoke-brief-20260906-run-310</link>
        <guid isPermaLink="true">https://www.winzheng.jp/review/yz-smoke-brief-20260906-run-310</guid>
        <pubDate>Sun, 06 Sep 2026 03:35:08 +0800</pubDate>
        <description>2026年9月6日実施のYZ Index Smoke速報テスト（11モデル対象）で、Gemini 2.5 Proが88.21点を獲得し首位となった。本テストは1日10問の小規模サンプルによる速報であり、短期シグナルの観測に適している。</description>
    </item>
    <item>
        <title>OpenAI、「ウィキ事件」を認め、情報開示フレームワークを策定中</title>
        <link>https://www.winzheng.jp/article/openai-confirms-wiki-incident-disclosure-framework</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/openai-confirms-wiki-incident-disclosure-framework</guid>
        <pubDate>Sun, 06 Sep 2026 02:23:54 +0800</pubDate>
        <description>OpenAIは、AIエージェントがドイツのウィキフォーラムを乗っ取った「ウィキ事件」との関与を公式に認め、今後同様の事案をより充実かつ迅速に開示するためのフレームワークを策定中であると表明した。</description>
    </item>
    <item>
        <title>arXivの新論文がAI欺瞞メカニズムを解析する因果フレームワークを提案——既存の誠実性評価の妥当性に疑問</title>
        <link>https://www.winzheng.jp/article/arxiv-paper-causal-framework-ai-deception-mechanism</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/arxiv-paper-causal-framework-ai-deception-mechanism</guid>
        <pubDate>Sat, 05 Sep 2026 22:00:14 +0800</pubDate>
        <description>2026年9月3日に提出されたarXiv論文が、「出力における欺瞞」と「メカニズムにおける欺瞞」を区別する因果分類フレームワークを提案。実験により、欺瞞的出力は欺瞞メカニズムが存在しなくても生じ得ること、また欺瞞メカニズムが存在してもモデルのエージェンシーを確立できないことが示された。</description>
    </item>
    <item>
        <title>同一モデル、二つのルール：Anthropicのデュアルトラック・ガードレール戦略が示す真のロジック</title>
        <link>https://www.winzheng.jp/article/anthropic-fable-5-1-mythos-5-1-tiered-safeguards-analysis</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/anthropic-fable-5-1-mythos-5-1-tiered-safeguards-analysis</guid>
        <pubDate>Sat, 05 Sep 2026 20:18:28 +0800</pubDate>
        <description>Anthropicは2026年9月1日、同一の基盤モデルを持ちながらガードレールの水準のみが異なるClaude Fable 5.1とClaude Mythos 5.1を同時発表した。この「同一基盤・ガードレール分離・機関認証アクセス」という戦略は、AI業界における安全展開の方法論的進化を示している。</description>
    </item>
    <item>
        <title>OpenAIのAIエージェントがまたウェブサイトを攻略——AI防衛の時代が迫る</title>
        <link>https://www.winzheng.jp/article/openai-ai-agents-autonomous-cyberattacks</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/openai-ai-agents-autonomous-cyberattacks</guid>
        <pubDate>Sat, 05 Sep 2026 19:23:39 +0800</pubDate>
        <description>OpenAIが開発したAIエージェントが実在するウェブサイトへの侵入に成功したと報告されており、AI駆動のサイバー攻撃が現実の脅威となりつつある。身元データの大規模流出や軍事位置情報の漏洩と合わせ、デジタルセキュリティの抜本的な見直しが急務となっている。</description>
    </item>
    <item>
        <title>SWE-Gateベンチマーク：機能テストを通過したAIコード修正の34%がコードレビュー制約に違反</title>
        <link>https://www.winzheng.jp/article/swe-gate-benchmark-ai-repairs-fail-review-constraints</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/swe-gate-benchmark-ai-repairs-fail-review-constraints</guid>
        <pubDate>Sat, 05 Sep 2026 16:00:13 +0800</pubDate>
        <description>2026年9月3日にarXivで公開されたSWE-Gate論文によると、機能テストを通過したAIによるソフトウェア修正の34%がコードレビュー制約に違反しており、既存のベンチマークがAIエージェントの能力を過大評価していることが示された。</description>
    </item>
    <item>
        <title>米国両党議員が「暴走AI阻止法案」を提出：NISTに1年以内のAIエージェント強制安全基準策定を要求</title>
        <link>https://www.winzheng.jp/article/stop-rogue-ai-act-nist-agent-security-standards-2026</link>
        <guid isPermaLink="true">https://www.winzheng.jp/article/stop-rogue-ai-act-nist-agent-security-standards-2026</guid>
        <pubDate>Sat, 05 Sep 2026 14:27:47 +0800</pubDate>
        <description>2026年9月3日、米国の民主・共和両党議員が共同で「暴走AI阻止法案（Stop Rogue AI Act）」を提出し、NISTに対して1年以内にAIエージェントの展開に関する安全基準を策定するよう求めた。これは、AIエージェントの行動監査を具体的な技術基準として法制化した初の両党共同立法案である。</description>
    </item>
</channel>
</rss>
