The Registerが社内関係者からのタレコミとして報道。IBM傘下のRed Hatが、これまでのAIツール積極推進の姿勢を転換し、R&D部門の開発者に対しAIトークン利用額を月300ドルまでに制限する通達を出したと伝えた。加えて、余ったトークン割当を他の開発者と共有することも禁止された。Red HatとGartnerの双方に取材を申し込んだが、記事執筆時点でいずれも回答はなかった。
セントルイス連銀・ヴァンダービルト大・ハーバード大の経済学者4名(Bick, Blandin, Deming, Schumacher)が、OpenAI/Anthropic/Microsoftのチャットログ分析ではなく、Real-Time Population Survey(RPS)データを使って労働者の実際の生成AI利用を計測した論文"What Work Does Generative AI Do?"を分析。チャットログ分類器は「文書の編集」のような汎用タスク記述に紐づけるため、O*NETデータベースの実際の職務内容と食い違い、AI企業側の推計は生成AIの職務関連度を過大評価していると指摘。
"For example, four out of five detailed occupations have adoption rates above 20 percent, but only one out of six occupations exceed 70 percent adoption."
Salesforceの副CFO Mike SpencerがDeutsche Bank Technology Conferenceで語った内容。約6ヶ月前からR&Dサイクルに Claude を投入し、どこまで開発が加速できるかを試す取り組みを行った結果、Claudeトークンへの支出がかさみ、それを補うため通期の営業利益率ガイダンスを引き上げなかったと説明。現在は「タスクに応じて適切なモデルを選ぶ」refinement modeに移行し、全タスクに最新モデルを使う必要はなく、OpenAI/Cursor/Claude/Grokなど複数ベンダーのコスト構造を比較しながら最適化を進めている。
Salesforce said its operating margin, according to accounting rules, would be 20.5 percent in its Q2 results (ending July 31), but its guidance for the full year is 20.1 percent.
"The IT advisory firm surveyed more than 1,300 leaders from organizations with more than $50 million in revenue between January and April of this year."
Financial Times が入手したAmazon社内資料に基づきTechRadarが報じた。著者情報と商品リストを紐付ける社内向けのClaude Sonnet導入プロジェクトが、当初予算を860%超過し総額180万ドルに膨張。この超過は約5ヶ月間社内で検知されないままだった。同じ資料は他の暴走事例も指摘しており、財務監査ツール開発プロジェクトが54.1万ドル、配送短縮を狙った物流プロジェクトが13.4万ドルの想定外コストをそれぞれ計上した。社内関係者は、サブスクリプション課金からトークン従量課金へ移行したことで、以前なら軽微だったミスが破滅的なコストに変わったと説明している。
Duolingoは2025年4月、CEOルイス・フォン・アンが「AIファースト」方針を発表し、社員の人事評価にAI活用度を組み込み、AIで代替できない業務のみ新規採用を認めるとした。この方針は社内外から強い反発を招き、ユーザーからはアプリ削除の表明も相次いだ。約1年後の2026年4月、フォン・アンはSilicon Valley Girlポッドキャストで方針転換を明言し、AI活用そのものを評価基準にするのをやめ業務成果で評価する方針へ戻すと説明した。
IBM Institute for Business Value (IBM IBV) が Salesforce顧客1,200人超を対象に調査した「State of Salesforce 2025-2026」レポート。AI/エージェント型AI(Agentforce等)導入の成果を、ROI達成率・全社展開の可否・頓挫率で定量化した。
"For 1,200+ Salesforce customers surveyed, only 33% of AI initiatives are meeting ROI targets. Even more concerning: 72% have failed to scale across business units, and 20% have stalled, failed outright, or been abandoned."
"McDonald's is ending its two-year-old test of drive-thru, automated order taking (AOT) that it has conducted with IBM and plans to remove the technology from the more than 100 restaurants that have been using it... the technology will be shut off in all restaurants currently testing it no later than July 26, 2024." / "there have been questions about whether that technology is ready for prime time, amid concerns about order accuracy."
MIT NANDA系の「GenAI Divide」調査(95%失敗/5%成功)を扱った記事で、canonical URL・og:title・articleAuthor(Jason Snyder)が候補URLと完全一致し、本文にも95%/5%/83%/90%等の数値が実際に登場することを確認した(前回却下は別記事への誤到達が原因で、今回のcurl取得では正しい記事が返った)。ただし本文中に "NANDA" という語自体は0件で、レポート名は "State of AI in Business 2025" とのみ記載。数値の一次ソースはMITレポートであり、Forbes記事はその解説記事(独立メディアによる紹介記事)という位置づけ。
"MIT’s report highlights a “shadow AI” economy, where employees in over 90% of firms use personal AI tools even when official pilots fail... employees quietly rely on personal AI to speed up claims processing, part of a pattern that MIT says is already saving companies $2 million to $10 million a year in external costs and cutting agency spend by 30%." / "The data is stark: Only 5% of custom GenAI tools survive the pilot-to-production cliff, while generic chatbots hit 83% adoption for trivial tasks but stall the moment workflows demand context and customization."
Robert Halfの調査で「調査対象企業の約29%が、AI導入を理由に人員削減した後、その従業員を再雇用した」と報じられている。
About 29% of companies surveyed had laid off workers after implementing AI only to rehire them, according to a study by Robert Half, a talent solutions and business consulting firm.
Forbes contributor article (TerDawn DeBoe, published 2026-05-21, updated 2026-07-09) citing third-party survey/research data: Robert Half survey and Forrester executive survey, plus a Time magazine report on the rehire pattern. Article describes companies cutting staff after announcing AI would take over a role, then rehiring the same staff 6-12 months later once the AI covered only part of the job duties. Gives customer service, marketing (copywriting), accounting, and sales as example function categories (no named companies).
示唆
Article's central claim: AI replaced tasks, not full jobs — it handled routine parts (answering, transaction processing, lead qualification) but failed at judgment, memory of past interactions, anomaly detection, and relationship-building, forcing partial rehires. Rehired workers often needed new skills to manage AI tools and were paid more, raising total cost. Small businesses face a thinner margin for error than large enterprises when the same substitution fails.
成果・金額自己申告
Robert Half: 29% of organizations that cut employees due to an AI-related reduction had already rehired into the positions they cut (article's actual figure — differs from the 32% in the task hint). Separately, Forrester: 55% of executive decision-makers who replaced employees with AI expect to regret the decision within 18 months. Article also notes a wage jump example: a $55,000/year role rehired at $75,000/year to also manage AI tools.
"Robert Half reported that a total of 29% of those organizations which cut employees due to an AI related reduction had already rehired into the positions they cut. According to Forrester, 55% of executive decision makers who replaced their employees using AI will regret the decision within 18 months."
"Orgvue, 39% of business leaders made employees redundant due to AI deployment. However, among that number, 55% admit wrong decisions about those redundancies were made." / "Meanwhile, 32% of U.S. hiring managers said they eliminated a role primarily due to AI and later rehired for the same or a similar position, according to data from Robert Half sent to CNBC."
IBM Researchのローハン・アローラ氏らが、AIモデルの障害対応能力を科学的・定量的に評価するオープンソース検証システム「ITBench」を開発。Kubernetes上で稼働し、Ansibleで障害を注入し、AIエージェントの診断・緩和能力を「Pass at K」「Mean Time to Resolution」等の指標で測定。初期実験ではLLM(Gemini 3 Pro等)にデータ収集〜解決策立案まで全て任せ、「ReAct」等の自律調査プロセスを用いたが、原因特定の成功率は約14%、問題解決に至る割合は約11%にとどまった。原因はコンテキストに混じるノイズ(無関係な情報)にLLMが固執し、巨大で複雑なKubernetesシステムで迷子になり不要な情報収集を繰り返してハルシネーションを起こすこと。改善策として、アプリケーションのトポロジー情報を用いてLLMの探索範囲をアラート発生サービス周辺に限定し、固定手順で状況共有する制約付きアプローチを導入した結果、3回試行(Pass at 3)での原因特定精度が約95%まで向上した。
MicrosoftとThe Hong Kong University of Science and Technology(HKUST)/清華大学系の研究チームが、プログラミング実行箇所の特定・ログ解析・相関関係の判定という一連の作業工程を対象に、実データベースの大規模言語モデル(Claude 3.5含む)テストを実施。各工程を通しで評価したところ、いずれかの段階で誤りを含めると原因特定の総合精度が11%以下まで落ち込むことを確認した。