Skip to content

Latest commit

 

History

History
258 lines (234 loc) · 46.8 KB

File metadata and controls

258 lines (234 loc) · 46.8 KB

Watchlist

Open threads that span multiple days — so nothing drops between editions.

Status legend: 🟢 confirmed/closed · 🟡 active/developing · 🔴 stalled · ⚪ rumor

Last updated: 2026-05-22


Funding & Valuations

Thread Status Last move Watching for
Anthropic $30–50B raise at up to ~$950B post 🟡 2026-05-18: still no term sheet signed (per BuildFastWithAI Monday digest); 3rd consecutive week characterized as "imminent" — likely slipping to align with post-June-15 metering data; possible close timed to Code w/ Claude Tokyo (June 5–6) Signed term sheet; final post-money; lock-up period; whether MGX / UK Sovereign / SoftBank Vision Fund 3 anchor
Anthropic ARR trajectory 🟡 $44B disclosed in May 8 investor meeting Monthly progression; first month of deceleration
Anthropic first profitable quarter 🟡 2026-05-21: reporting projects Q2 2026 revenue $10.9B (>2×) and **$559M operating profit — ~2 yrs ahead of internal plan**, despite $15B/yr Colossus bill Whether profitability holds once the full $1.25B/mo compute cost ramps; impact on the $30–50B raise terms / IPO timing
Anthropic > OpenAI on US business adoption 🟢 2026-05-14 Ramp AI Index: Anthropic 34.4% vs OpenAI 32.3% — first crossover ever Revenue crossover (Anthropic leads seats, OpenAI leads dollars); whether the 2.1pt lead holds
OpenAI "The Development Company" enterprise JV 🟡 Raised $4B from 19 investors at $10B valuation First enterprise customer logos; how it competes with Anthropic's PE-deployment JV
Wispr Flow ~$260M / ~$2B 🟡 In talks per Bloomberg May 12; Menlo Ventures leading Round close; Voice OS product expansion
Judgment Labs $32M Seed + Series A 🟢 Closed May 12 — Lightspeed leading both rounds Hiring volume; first major customer logos
Anthropic Oct 2026 IPO path 🟡 Active consideration alongside the $50B round S-1 filing; whether it's primary or secondary-heavy
OpenAI IPO path (confidential S-1) 🟡 2026-05-22: filing a confidential S-1 with the SEC as early as today (Goldman + Morgan Stanley); targeting Sept 2026 listing at ~$852B–$1T; financials stay private until ~15d pre-roadshow; unblocked after Musk lost his lawsuit Public S-1 (revenue mix: API vs ChatGPT vs ads vs enterprise; Microsoft terms; ad-revenue disclosure); roadshow window; first-day pop → first-earnings arc
Anthropic IPO path (October 2026 target) 🟡 2026-05-22: reported targeting October 2026 listing, on the back of projected first profitable quarter (Q2 ~$10.9B rev, ~$559M op profit) Whether the October path firms up; S-1 filing; how the $1.25B/mo Colossus bill reads in a prospectus
The IPO wave as an asset-class shift (SpaceX + OpenAI + Anthropic) 🟡 NEW 2026-05-22: three frontier-adjacent giants potentially public inside ~12 mo = frontier AI becomes a public-market asset class; changes liquidity, secondary comp, employer-risk First prints; whether public markets underwrite frontier-AI economics; founder-recycling from liquidity events
OpenAI $25B ARR + IPO path 🟢 SUPERSEDED Q1 2026 IPO chatter → now concrete (see confidential S-1 row above) (folded into the OpenAI IPO-path row)
OpenAI ad revenue (ChatGPT Ads Manager) 🟡 2026-05-21: self-serve Ads Manager (CPC, no minimum spend, agency + Adobe/Criteo partners) live; target $2.5B this year → $100B/yr by 2030; ads to Free/Go tiers — vs Anthropic's ad-free pledge Whether ad revenue materializes at target; whether Anthropic holds ad-free; agent-mediated-commerce / attribution as an emerging wedge
Cognition (Devin) $25B raise 🟡 Hundreds of millions pending; Founders Fund leading Round close; final valuation; new enterprise logos
Sierra $15.8B Series E 🟢 Closed early May Hiring volume / TAM signals over next 60 days
Moonshot AI $20B valuation 🟢 Closed early May Kimi K2.7 release; OpenRouter share
Q1 2026 venture data ($300B / 80% AI) 🟢 Crunchbase reported Q2 2026 print — will the AI share hold above 70%?
Browser-agent infra category ($2B+ rounds) 🟡 Parallel Web closed at $2B Next $2B+ round in the category
Mistral capital event (EU sovereignty catalyst) EU vs. Mythos escalation accelerates this Round $5B+ announced; or xAI partnership formalized
Anthropic acquires Stainless (≥$300M) 🟡 2026-05-12 → 2026-05-15: advanced talks reported (The Information, multiple confirms) Deal close; OpenAI/Google response on SDK toolchain ownership
Chapter Medicare-AI Series E ($100M, Generation IM) 🟢 Closed May 2026 Hiring volume; whether Medicaid-equivalent player gets funded next
GridCARE $64M Series A (oversubscribed, power-acceleration) 🟢 Closed week of May 12; significant step-up from prior round First adjacent round in interconnection-queue or on-site-gen tooling; NVIDIA strategic invest?
Exaforce $125M Series B (agentic SOC) 🟢 NEW 2026-05-22 (closed ~May 12): HarbourVest/Peak XV/Mayfield/Khosla/Seligman/AICONIC; total $200M (1 yr after $75M A); real-time security knowledge graph + agents ("Exabots") + MDR; ~10× faster investigations, fewer tokens; ~130 employees Next $100M+ AI-SOC round; whether the EO's cyber-clearinghouse formalizes a govt buyer; agentic-SOC hiring volume
Sprouts.ai $9M Pre-Series A (Revenue Agents, B2B GTM) 🟢 Closed May 15 (TGV + Accel) Hiring volume; whether the GTM-agent category gets another mega-round to anchor it
Nectar Social $30M Series A (agentic marketing OS) 🟢 Closed May 14 Customer logos; whether the consolidation pitch lands with mid-market marketers
Multiverse $70M (AI-skills workforce training) 🟢 Closed May 14 Enterprise customer expansion; pricing model
equipifi $34M Series B (BNPL for community banks) 🟢 Closed May 14 AI-product roadmap; partner-FI count
Isomorphic Labs $2.1B Series B (closed May 12) 🟢 2026-05-19 confirmed details: Thrive Capital lead; Alphabet/GV + CapitalG + MGX + Temasek + UK Sovereign AI Fund; ~$2.6B total capital; CEO Demis Hassabis; IsoDDE drug-design engine; partners Novartis + Lilly + J&J; London + Cambridge MA + Lausanne London / Cambridge MA / Lausanne hiring waves inside 60 days; whether four-corner template (Lab + VC + Sovereign + Industry) gets replicated in biotech/materials/defense in Q3
Sierra $950M at $15B (GV + Tiger Global lead) 🟢 2026-05-19 confirmed: $15B (not $15.8B) per Crunchbase; Bret Taylor + Clay Bavor; 3-year-old company; largest CX-agent valuation on record Sierra Customer Engineering / Solutions hiring volume; vertical-CX competitors raising in next 60 days; whether GV makes parallel agent-company bets
Parallel Web Systems $100M (Sequoia lead, total $230M) 🟢 2026-05-19: Parag Agrawal (ex-Twitter CEO); AI agent search/research infrastructure; second mega-round in the agent-search-infra category in 30 days Parallel customer logos; agent-identity / agent-friendly publisher contracts as adjacent unfunded layers
Runware $50M Series A (total $66M) 🟢 2026-05-19: "Sonic Inference Engine"; plans to deploy 2M+ Hugging Face models by EOY 2026 Whether long-tail-open-weights infrastructure category gets a $200M+ follow-on round; competitive responses from Replicate / Together / Modal / Fireworks
Oboe $16M (personalized course generation) 🟢 2026-05-19: Consumer EdTech-AI category Whether the B2B-equivalent (auto-generated SMB onboarding / compliance) is funded next
OpenAI Deployment Company ($4B+ initial capital) + Tomoro acquisition (~150 FDEs) 🟢 NEW 2026-05-19: Launched May 11–12; majority-owned subsidiary; TPG lead + 19-investor consortium (Advent, Bain Capital, Brookfield co-leads); Tomoro acquisition closes "in coming months"; Tomoro brings ~150 FDEs + offices in London/Edinburgh/Manchester/Singapore/Sydney/Melbourne; customers Tesco/Virgin Atlantic/Supercell Tomoro acquisition close date; first OpenAI Deployment Company customer logo announcements; comp-band setting for senior FDEs (likely $300–500K base, $700K+ TC at staff); whether Anthropic Solutions matches or differentiates

Models & Capability

Thread Status Last move Watching for
GPT-5.6 / next OpenAI flagship Leak chatter on Latent Space; nothing official Official announcement; Realtime-3 voice
Claude Opus 4.8 / next Anthropic flagship No leaks yet Anthropic dev keynote (rumored June)
Gemini 3.5 (Flash shipped; Pro pending) 🟢 2026-05-19 RESOLVED: Gemini 3.5 Flash GA same-day — $1.50/1M in · $9/1M out · $0.15 cached · 1,048,576 in / 65,536 out ctx · text+image+audio+video in · "within 2 pts of Anthropic flagship at ⅓ price" · already GA in GitHub Copilot. Gemini 3.5 Pro internal-only, ships June 2026 Gemini 3.5 Pro June launch + pricing; whether Claude/OpenAI cut prices in response; Flash adoption curve in agent stacks
WebMCP (open web standard for agents) 🟡 NEW 2026-05-19: Google proposed WebMCP at I/O — sites expose callable JS/HTML-form tools; browser agents call instead of scrape. Built on Anthropic MCP lineage. Experimental origin trial in Chrome 149; Gemini-in-Chrome support "coming soon" Chrome 149 flag availability date; whether OpenAI adopts MCP-shaped web tooling; first WebMCP-native infra startups (60–90d lag expected); W3C standardization path
Gemini 3.5 Flash price war ($1.50/1M) 🟡 NEW 2026-05-19: VentureBeat — can "slash enterprise AI costs $1B+/yr"; reframes frontier as cheapest-good-enough + best rails Claude / OpenAI price response; impact on "cheap-inference" startups (Runware et al.); whether cost-aware routing becomes a standard interview topic
Antigravity 2.0 + Managed Agents (Gemini API) + ADK 2.0 🟡 NEW 2026-05-19: standalone agent-first desktop app + CLI + SDK; "one API call → sandboxed agent (reason/tools/code in isolated Linux)" — near-verbatim Anthropic Managed Agents; Chrome DevTools-for-agents supports 20+ non-Google agents Google Cloud Agent/Antigravity Solutions hiring; Gemini Enterprise Agent Platform pricing detail; whether the runtime primitive fully commoditizes
Gemini Spark + AI Ultra ($100/mo proactive consumer agent) 🟡 NEW 2026-05-19: "24/7 AI agent" in new $100/mo tier; matches Anthropic Max-5x price point Spark adoption; whether proactive-agent trust/safety incidents surface (ties to IPI thread); OpenAI/Anthropic consumer-agent responses
Nemotron 4 (NVIDIA open) Nemotron 3 Nano Omni shipped May Whether NVIDIA goes annual on Nemotron
DeepSeek V5 V4 still being discounted Summer launch likely
Mistral open frontier release Quiet quarter EU sovereignty narrative catalyst
Apple "Extensions" SDK 🟡 iOS 27 reveal expected WWDC June 9 Official SDK + featured Extensions list
xAI voice stack (STT + TTS + Imagine Quality) 🟢 Standalone APIs shipped GA week of May 11 Mistral partnership formalization; speech latency benchmarks
Googlebook / Aluminium OS (Google desktop platform) 🟡 Codename confirmed + rebranded "Googlebook" May 12–13; OEMs locked (Acer/Asus/Dell/HP/Lenovo); Gemini as OS layer Developer SDK at I/O May 19; OEM hardware shipping dates; "Magic Pointer" agent surface
Claude Code share of GitHub commits 🟡 SemiAnalysis: ~4% of all public commits (~135K/day), double the prior month Whether the 20%-by-end-2026 projection tracks; reliability/outage complaints at scale
Android XR glasses (gen 2) 🟡 Previewed May 12 Final hardware partners; consumer launch date
Claude for Legal (12 plugins + 20+ MCP connectors) 🟢 Launched May 12 — first major Anthropic vertical Healthcare / Finance / etc verticals next; partner volume growth
PwC × Anthropic: 30K trained → 364K global; Claude-native Finance practice 🟡 2026-05-14 announcement; first Big-4 single-vendor mass-cert commitment Deloitte / Accenture / EY counter-commitments (90 day window); first F500 deal won by the Finance practice
Google I/O 2026 (May 19–20) 🟡 TODAY — Day-1 keynote 10 AM PT. Final pre-keynote consensus narrowed to: Gemini 3.5 (most leaks), Gemma 4 open-weights, Gemini Omni unified video/image/audio, Gemini Spark / Remy proactive agent, Android XR Gen 2 (Samsung + Warby Parker + Gentle Monster all named), Aluminium OS / Googlebook formal name + OEM ship windows (Acer/Asus/Dell/HP/Lenovo), Android 17 SDK with system-level agent hooks, Vertex AI Agent Platform pricing (current rate $0.0864/vCPU-hr + $0.25 per 1K events; Gemini 3 endpoint pricing live July 1) Sundar's first-8-min consumer-vs-enterprise framing; Vertex Agent Platform pricing vs Claude $20/$100/$200 metered tier (post-June 15); first independent Gemma 4 fine-tune to cross 1K stars (likely Thursday); Day-2 developer breakouts May 20
Meta Avocado / Mango (next flagship LLMs) 🟡 Delays confirmed May 12; closed-source pivot at C-suite H2 2026 launch date; whether any portion ships open-weights
Claude for Small Business (QuickBooks/PayPal/HubSpot/Canva/DocuSign/Workspace/MS365) 🟢 Launched May 13; free 10-city in-person tour kicked off May 14 Adoption velocity in SMB segment (Anthropic's first <500-employee push); partner-app integration depth
Anthropic Agent SDK metering split (effective June 15) 🟡 2026-05-18: confirmed via TECHSY that the credit doesn't auto-activate — manual toggle required in account settings (do this tonight). Announced; credit pool by tier ($20/$100/$200); programmatic billed at full API list rate June 8 email with exact allocations; how many programmatic-Claude users miss the toggle and get blocked on June 15; community workarounds (router SDKs, prompt caching, batching); downstream vendor responses (Zed already adapting)
OpenAI ChatGPT Personal Finance (Plaid, 12K+ FIs) 🟢 Launched May 15, Pro-only US preview Intuit integration timing; Plus rollout; non-US clones
Gemini "Omni" unified video/image/audio model 🟡 Sample clips now public from at least one Gemini Pro tester (May 16–17); ~10 sec/clip cap suggested, daily quota burnable in 2 prompts Official reveal May 19 I/O keynote
OpenAI Codex in ChatGPT mobile (iOS + Android) 🟢 Shipped May 14 to all plans incl. Free/Go; macOS desktop-app prereq; ~4M weekly Codex users disclosed Windows desktop support timing; whether Anthropic ships Claude Code mobile parity within 90 days
Anthropic × Gates Foundation $200M / 4-yr partnership 🟢 Announced May 14, full Sunday cycle May 17 — global health (polio, HPV, eclampsia/preeclampsia), education (K-12 + sub-Saharan Africa / India literacy), economic mobility (smallholder farming) First Anthropic JD posts for "Applied AI — Mission Programs / Solutions Engineer — Health / Education" within 30 days; OpenAI / DeepMind parity mission deals
Code w/ Claude London (May 19) + Tokyo (June 10) 🟢 2026-05-19 CORRECTION: Anthropic's official site lists London = May 19 (TODAY, same calendar day as I/O 2026) and Tokyo = June 10. Previously logged as London May 20–21 and Tokyo June 5–6. The same-day collision with I/O is sharper counter-programming than the +36-hour framing. Day-1 livestream; keynote panel Ami Vora · Boris Cherny · Angela Jiang; customer presenters Asana · Cursor · GitHub · Replit · Vercel New SDK feature announcements timed to land 3–4 hours after Sundar; whether Anthropic reveals a "Vertex-Agent-Platform response"; which customer announces a deeper Claude integration; whether Tokyo announces APAC customer presenters by end of May
Isomorphic Labs $2.1B Series B 🟢 NEW Closed May 12 — Thrive lead, + Alphabet/GV + MGX + Temasek + CapitalG + UK Sovereign AI Fund; total capital base $2.6B; first "Lab+VC+Sovereign+Industry" four-corner template Whether the four-corner template is replicated in another vertical AI round inside 60 days (biotech / materials / energy / defense); London + Cambridge MA + Lausanne hiring waves
Mustafa Suleyman 18-month white-collar automation forecast 🟡 NEW Multiple May 2026 reiterations (Fortune, Tom's Hardware, Irish Examiner, AzazTV) — names accounting, legal, marketing, project management as first targets; opposing data from Thomson Reuters 2025 survey shows augmentation > displacement Whether Sam Altman + Sundar Pichai publicly converge on the same 12–18 month timeline at I/O / future events; whether Copilot-attach economics show the predicted lift in Microsoft Q2 earnings (Aug 2026)
OpenAI Sweetpea / Jony Ive device 🟡 H2 2026 target reaffirmed (Chris Lehane at Axios House Davos Jan 2026); Foxconn manufacturing; screenless behind-the-ear wearable + glasses + voice-recorder exploration Official reveal date; first hardware impressions; pricing; OS integration story (does it embed ChatGPT mobile + Codex Mobile?)
Anthropic ad-free policy commitment 🟢 NEW 2026-05-19: Anthropic position post stating Claude remains ad-free; advertising incentives "incompatible with a genuinely helpful AI assistant"; carves out competitive moat vs Google's consumer Gemini Spark / ad-supported tiers Whether OpenAI / Google publish parity commitments; whether Anthropic ad-free policy applies to all enterprise integrations or only to Claude.ai consumer surface
Workday Foundation × Anthropic Solopreneurship Accelerator (15 slots) 🟢 NEW 2026-05-19: 15-founder cohort; seed funding + Claude credits + AI-first entrepreneurship curriculum Application deadline; first-cohort wedge selection (indicates which Anthropic verticals the accelerator favors); whether Workday + Anthropic expand to a 50-slot cohort in 2027

Policy & Geopolitics

Thread Status Last move Watching for
EU AI Act enforcement window opens August 2026 🟡 Article 51 enforcement window starts in 3 months First Article 51 invocation; first fine notification
US CAISI pre-deployment reviews 🟢 MS, Google, xAI, OpenAI, Anthropic all signed Will EU build a mirror; first non-public test result
Anthropic Mythos EU access standoff 🟡 Spain Minister Cuerpo publicly cited Article 51 (May 5–11) Formal EU enforcement action; counter-narrative from Anthropic; OpenAI takes EU market share
OpenAI GPT-5.5-Cyber EU access deal 🟢 Confirmed agreement to share with EU pre-deployment body (CNBC May 11) Formal review timeline; how it differs from Mythos terms
Pentagon: 8 AI vendors selected, Anthropic excluded 🟢 Confirmed early May Anthropic appeal or alternative DoD path
First AI-built zero-day in the wild (Google Threat Intel) 🟢 Confirmed by Bloomberg + Google May 11–12; mass-exploit campaign averted Whether NIST issues updated AI-risk guidance within 60 days; CISO mandate response
Trump AI Action Plan execution 🟡 CAISI partnerships are first deliverable Export-control updates; sovereignty mandates
Trump AI/cybersecurity executive order (pre-release model review) 🟡 STALLED 2026-05-22: POSTPONED — the planned signing was pulled; Trump "didn't like certain aspects" and "hates regulation" ("I don't want to get in the way of [US AI] leading"). Draft survives: voluntary 90-day pre-release frontier review + Treasury-led cybersecurity "clearinghouse" (find/fix vulns in unreleased models); negotiated w/ Nat'l Cyber Director Sean Cairncross + OpenAI/Anthropic/Reflection AI. 2026-05-21: had been "signing as soon as Thursday w/ CEO ceremony"; labs lobbying 14 vs 90 days The re-scheduled signing date (if any); whether the 90-day frontier-review half survives a redraft or only the cyber-clearinghouse does; the (now-delayed) pre-deployment-eval / AI-assurance job market (see 2026-05-22/01 §1)
US–China AI safety protocol 🟡 2026-05-14: Bessent confirms formal talks launched at Trump–Xi Beijing summit; goal = keep frontier models from non-state actors Published provisions; weight-custody / API-KYC requirements; first concrete compliance surface
Microsoft–OpenAI partnership amendment 🟢 Non-exclusive license confirmed May; OpenAI can serve from any cloud Whether OpenAI signs major non-Azure cloud deal next

Compute & Infrastructure

Thread Status Last move Watching for
Anthropic–Google $200B compute deal (5 yrs) 🟢 Confirmed early May Quarterly utilization disclosures
Anthropic rents all of Colossus 1 from xAI/SpaceX 🟢 CONTRACTUAL 2026-05-21: terms now disclosed in SpaceX's S-1$1.25B/month through May 2029 (~$15B/yr, $40B+ total), discounted first 2 months during ramp; entire Colossus 1 = 300MW, 220K+ GPUs (H100/H200/GB200); Musk frames it as "AI compute as a service at scale" Whether Colossus 2/3 deal materializes; whether the $15B/yr bill survives Anthropic's profitability projection; other labs disclosing comparable compute liabilities in filings
NVIDIA $40B+ in AI equity bets 🟡 IREN $2.1B, Corning $3.2B confirmed Next major equity bet; antitrust angle
US grid load 5–7% AI by 2027 🟡 Air Street State of AI projection Data-center permitting wins/losses

Jobs & Hiring Signals

Thread Status Last move Watching for
Karpathy → Anthropic (pre-training automation) 🟡 NEW 2026-05-22 (announced May 19, started this week): Andrej Karpathy (OpenAI founding member → Tesla → OpenAI → Eureka Labs) joins Anthropic's pre-training team, launching a new group using Claude to accelerate pre-training research What the team ships; whether it's the production face of the PostTrainBench / recursive-self-improvement thread; downstream "AI-does-AI-R&D" tooling startups; talent-market follow-on (Anthropic queue thickening)
CAIO adoption 76% (IBM) 🟢 Report released May 11; broader coverage May 12 LinkedIn job-posting volume for CAIO + reports-to-CAIO roles
CS new-grad market bifurcation (Q1 data — revised) 🟢 Revised May 13: 78,557 Q1 tech layoffs / 47.9% AI-attributed · generic SWE -40% · MLE +41.8% YoY Q2 print; whether MLE growth holds at >25% YoY
Meta May 20 layoffs (8,000 ≈ 10% workforce) 🟢 EXECUTING 2026-05-20: notifications began Wed — ~8K notified (≈10%) + 6K canceled open reqs = ~14K impact; Singapore first (4 AM local) → UK → US; ~7,000 redirected into new AI teams: Applied AI Engineering / Agent Transformation Accelerator XFN / Central Analytics (CPO Janelle Gale); AI infra spend cited up to $145B; more company-wide cuts planned H2 2026 (Reuters); confirmed NPR/CNN/CNBC/Reuters; severance 16 wks + 2 wks/yr + 18 mo health Outreach window Thu 5/21 8 AM PT — split pool (b) displaced SDE/ads/ops vs (c) redirected-to-AI; Meta-alumni startup formations 60–90 days; H2-cut timing; Meta Q2 GAAP operating-income as the "AI-capex-substitutes-headcount" empirical test; whether MS/Salesforce/Oracle copy the redirect-not-just-cut move
Microsoft voluntary buyouts (~8,750 potential) 🟡 Confirmed ~7% US employees eligible Where ex-Microsoft AI/Copilot talent lands; small-AI-startup product orgs that absorb them
GM AI-skills-driven IT layoffs 🟢 Confirmed May 11 — "hundreds" cut to fund AI hires Atlassian model replication across legacy enterprises
Atlassian barbell (cut + 800 AI hires) 🟢 Confirmed; cleanest "AI-reshape" example Which large enterprises follow with the same simultaneous-cut-and-hire pattern
Frontier-lab new-grad MLE comp $400–600K 🟡 Anthropic-driven Whether Google/Meta/OpenAI publicly match in fall 2026 cycle
JPM 400 AI engineers, GS 200, MS 150, Citi 250 🟡 Confirmed throughout May When NYC AI-eng demand surpasses SF
FDE roles industry-wide 🟡 2026-05-19 MAJOR UPDATE: OpenAI Deployment Company launches with $4B + 150 Tomoro FDEs from Day 1 → makes OpenAI the best-funded FDE org in market. Updated landscape: OpenAI 150 → 500+ (Q3 projection) · Anthropic Solutions ~200 → 400+ · Palantir 700+ · Google Cloud 59 → 200+ (Vertex Agent Platform launches today) · Big-4 each 1,000+. Senior base $215–310K (Anthropic) vs likely $300–500K (OpenAI Deployment Co); TC at staff $700K+ at TPG-led OpenAI org Tomoro acquisition close date; comp-band differential between OpenAI Deployment Co vs Anthropic Solutions; Vertex Agent Platform FDE postings count by May 25; whether Big-4 announce parity org structures in Q3
Cognition (Devin) hiring wave 🟡 Round pending close; enterprise revenue 80× YoY Founding-engineer / deployment-engineer JD volume
Defense AI (Scout, Anduril, Shield, Saronic, Helsing) 🟡 Scout $100M Series A · Helsing $1.2B strategic Clearance-required role volume
Cloudflare layoffs (1,100) 🟢 Confirmed week of 05-09 Whether other "AI-replaced" layoffs in the same form factor follow
Cisco ~4,000 cuts + $9B AI-infra order guidance 🟢 Confirmed 2026-05-13; notifications start May 14; AI-order guidance raised $5B→$9B Where AI-infra hiring concentrates; whether other infra vendors print the same barbell
OpenAI Residency 2026 cohort 🟢 Applications open · $18,333/mo · career-changer focus Application deadlines · acceptance rates

Research Threads to Track

Thread Status Last move Watching for
AI does net-new mathematics (OpenAI Erdős result) 🟡 2026-05-21: an OpenAI general-purpose reasoning model disproved a central conjecture in the planar unit-distance problem (Erdős 1946) — infinite family of better constructions, proof via algebraic number theory; verified by Noga Alon + Thomas Bloom Reproducibility / generality — can the same model do this across problems, or needle-in-haystack? Independent replication; whether other labs publish comparable "general model → novel proof" results; live-benchmark response (LemmaBench)
Live / contamination-resistant benchmarks (LemmaBench, RepoReason) 🟡 2026-05-21: arXiv wave of continuously-refreshed, memorization-resistant evals (LemmaBench 2602.24173 live math; RepoReason 2601.03731 repo-level; PostTrainBench 2603.08640) Adoption by labs / the EO pre-release review process; whether eval-as-a-service startups raise on it
Real-tool agent benchmarks (MCP-Atlas, Tool Decathlon) 🟡 NEW 2026-05-22: eval shifts from mocks to real tools — MCP-Atlas (Scale, arXiv 2602.00933; real MCP servers, agent must discover tools) + Tool Decathlon/Toolathlon (ICLR 2026, arXiv 2510.25726; 32 apps / 604 tools, execution-based eval) Whether labs report MCP-Atlas/Toolathlon in release notes; eval-against-real-tools as a recurring-revenue startup wedge; whether these become standard FDE-interview vocabulary
Agentic Reasoning survey — 3-layer taxonomy 🟢 NEW 2026-05-22: arXiv 2601.12538 — foundational (plan/tool/search) → self-evolving (feedback/adaptation) → collective (multi-agent); the connective vocabulary under the week's benchmark + talent moves Citation velocity; whether the taxonomy terms appear in job-post requirements / lab posts
Anthropic Dreaming 🟡 Research preview May 7 Generalization beyond coding/finance/legal
Agent Reliability framework (arXiv 2602.16666) 🟢 v2 active; 12 metrics across 4 dimensions; reliability decoupling thesis Replications; whether labs internally adopt the metrics
Outcome-Driven Constraint Violations benchmark (2512.20798) 🟢 Published; complements Constraint Decay (capability vs alignment failure modes) Independent benchmarks; vendor-by-vendor replication
HF trending: GenericAgent / rStar / HyperEyes / AutoTTS / DeepCode 🟢 Trending May W20 Whether DeepCode-style doc-to-code becomes a startup category
Multi-agent vs single-agent under matched compute 🟡 Stanford finding Replications; counter-papers
Mamba/hybrid open frontier (Nemotron) 🟡 Nano Omni shipped Next hybrid open model at scale
On-policy distillation sweep (OPD survey + SDPO + OPSD) 🟡 First comprehensive survey published April · SDPO + OPSD reproducible code Whether frontier labs adopt at training time
Constraint Decay (arXiv 2605.06445) 🟡 Submitted May 7 — replicable failure mode in coding agents Independent replications · mitigations from coding-agent vendors
Structured Distillation for Personalized Agent Memory 🟡 11× compression with retrieval preservation Adoption inside Cognition / Sierra / Decagon stacks
Mem0 / EverMemOS memory architectures 🟡 First production deployments When agent memory becomes standard
Externalization in LLM Agents (survey) 🟢 Unified review of memory, skills, protocols, harness eng Set the field's vocabulary; citation velocity
Calibration / abstention ("Answer, Refuse, or Guess?") 🟡 Appier research published — LLMs miscalibrated in both directions on risk Replications; whether abstention layers become standard in vertical agents
Agent memory maintenance (STALE / SAGE / survey 2603.07670) 🟡 Memory-validity + graph-memory work trending May W19 Whether staleness detection becomes a funded infra category vs Mem0/EverMemOS
Neuro-symbolic efficiency (~100× energy, ICRA Vienna) Presented May 2026 — task-specific result Replication outside robotics; whether it enables on-device agents
Attractor Models — fixed-point latent reasoning 🟡 arXiv submission May 12 — frames latent refinement as fixed-point problem; memory-efficient Replications outside the original lab; whether commercial APIs absorb the pattern
"Many Faces of On-Policy Distillation" (unified taxonomy) 🟡 arXiv May 11 — capability-gap sweet spot, diversity scaffolding, reward-landscape shape Whether labs cite this as the canonical OPD reference; production OPD pipelines that follow the prescription
AI-skill wage premium 25% → 56% in 12 months 🟢 Lightcast / Dice / IEEE Spectrum spring 2026 reads Whether premium holds above 40% through Q4 2026; non-AI tech wage compression continues
Cattle Trade multi-agent benchmark (arXiv 2605.14537) 🟢 Published May 14 — bluffing/bidding/bargaining under imperfect info Citation velocity; whether labs report scores in next flagship releases
Successor-Representation Spectrum for LLM topologies (arXiv 2605.11453) 🟢 Published May 12 — principled topology selection for multi-agent systems Adoption in multi-agent debugging tooling; integration into Sierra/Cognition/CrewAI
Topology-not-alignment safety position paper 🟡 Position paper week of May 12 (Bajaj et al.) Whether EU policymakers cite the framing; whether it shapes Article 51 enforcement criteria
DyTopo — Dynamic Agent Communication Topology 🟢 NEW Trending May 17 — graph rewires per round via semantic matching; complements Successor-Representation Spectrum Sierra/Cognition/Decagon production adoption; citation velocity 60-day check
AIRS-Bench — 20-task science-agent benchmark 🟢 NEW Trending May 17 — first reproducible science-agent benchmark; frontier scores 17%–34% end-to-end Whether labs report AIRS-Bench in flagship release notes; vertical AIRS-Bench equivalents (clinical / biotech / materials)
TrajAD — Runtime trajectory verifier w/ precise rollback 🟢 NEW Trending May 17 — Haiku-size verifier for Opus-size main agents; ~10× verifier-to-agent ratio Production-vendor integrations from Anthropic / Sierra / Cognition; false-positive-rate field reports
"Agentic AI Orchestration Should Be Bayes-consistent" position paper 🟡 NEW Position paper May 17 trending Whether frontier-lab posts cite the framing; whether "Bayes-consistent" appears in job-post requirements within 90 days
Lifting Traces to Logic (programmatic skill induction) 🟡 NEW arXiv May 17 — neuro-symbolic skill abstraction from agent traces; portable across providers Whether commercial agent memory systems (Mem0/EverMemOS) absorb the technique
CHAL — Council of Hierarchical Agentic Language Models (arXiv 2605.12718) 🟢 NEW arXiv May 12, trending May 18 — hierarchical multi-agent dialectic, defeasible argumentation as belief optimization; 3–8pp improvement on ambiguous-evidence tasks vs monolithic CoT, at 3–7× token cost Citation velocity by mid-July; production adoption by Sierra / Decagon / Cognition; whether the task-conditional multi-agent framing replaces the 2025 "single agent under matched compute wins" framing
MemReread — Memory-Guided Long-Context Rereading (arXiv 2605.10268) 🟢 NEW arXiv May 11, trending May 18 — streaming rereading guided by structured working memory; alternative to RAG for long-document tasks Whether the memory-architecture subfield converges on read-first vs retrieve-first; production adoption signals from Mem0 / EverMemOS / Cognee
ARIS — Cross-Model Adversarial-Collaboration Research Harness 🟢 NEW Trending May 18 — open-source counter to Anthropic Dreaming + Karpathy autoresearch; multi-vendor adversarial roles for verification Replication of Anthropic Dreaming results in public; whether commercial "AI scientist SaaS" products ship on top of ARIS
Storage Is Not Memory — Retrieval-Centered Agent Recall 🟡 NEW Trending May 18 — proposes storage / recall / memory trichotomy; reframes the "agent memory" product category Whether the vocabulary is adopted in product copy (Mem0, EverMemOS, LangMem, Cognee); citation velocity 60 days
Multimodal Procedural Knowledge — Skill Cards for Visual Agents (SJTU) 🟡 NEW Trending May 14, deeper May 18 — text + state-cards + visual-keyframes as portable skill structure for GUI / Computer-Use / robotics agents First commercial vendor to ship "portable skill cards" as a feature; whether a skill-card marketplace category forms
AIRS-Bench (20-task science-agent benchmark) 🟢 2026-05-19 update: Frontier scores 17–34% end-to-end; subset-task scores 50–70%; gap empirically anchors the multi-agent-coordination thesis (CHAL / DyTopo / Successor-Representation) Whether labs cite AIRS-Bench scores in I/O / Code w/ Claude keynote slides today; first community-submitted Claude / Gemini AIRS-Bench attempt published on GitHub
JADE — Expert-Grounded Dynamic Evaluation (per-claim decomposition) 🟢 NEW 2026-05-19: Per-claim factuality eval methodology — decompose agent output, check each claim against expert KB; foundational for vertical-AI eval (Claude for Legal, OpenAI personal finance) Whether Judgment Labs / Anthropic / OpenAI ship JADE-style APIs in production; first vertical-AI vendor to publish JADE-style scores on their own product
TrajAD — Runtime Trajectory Verifier with Precise Rollback (deeper) 🟢 2026-05-19: Haiku-size verifier monitors Opus-size main agent; ~10× verifier-to-agent ratio is cost-efficient enough to ship in production; rollback semantics are the missing primitive in current agent frameworks First open-source TrajAD implementation cross 1K stars; Anthropic / Sierra / Cognition production integrations
AgentScope distributed multi-agent improvements 🟢 NEW 2026-05-19: Distributed orchestration mechanisms + flexible environments + user-friendly tooling for large-scale multi-agent simulation Whether AgentScope becomes the canonical multi-agent reproducibility platform; citation velocity 60-day check
Indirect prompt injection (IPI) in the wild — Google threat report 🟢 NEW 2026-05-20: First large-scale empirical measurement — +32% relative growth in malicious IPI (Nov 2025→Feb 2026); real PayPal-transaction payloads hidden invisibly in HTML targeting payment-capable agents; sophistication still low (window to harden open); recommended defense = cheap dual-model "sanitiser" in front of privileged agent Next-quarter IPI growth number (leading indicator of agent deployment); whether dual-model sanitiser becomes a standard FDE-deployment requirement; first IPI-defense startup/feature; convergence with TrajAD/JADE (one supervisory primitive)
Convergence thesis — cheap guard model supervising privileged agent 🟡 NEW 2026-05-20: TrajAD (verification) + JADE (per-claim eval) + Google IPI (injection defense) all land on the same primitive; economically unlocked by Gemini 3.5 Flash @ $1.50/1M making always-on guard models ~free Whether labs ship "supervisor model" as a first-class API feature; whether the framing appears in job-post requirements within 90 days
DyTopo / RuleSmith / CommCP — multi-agent robustness wave 🟡 NEW 2026-05-20: DyTopo (per-round topology rewiring), RuleSmith (self-play + Bayesian opt for rule-balancing), CommCP (conformal prediction over inter-agent messages) — classical statistical rigor re-entering the agent stack Whether conformal prediction becomes standard in multi-agent coordination; production adoption; citation velocity

Personal Threads (For You)

Thread Status Watching for
Your demo project (overnight autoresearch run) Push to GitHub by May 18
Apply to 1 FDE role Submit by Sunday
Build 3-provider router Ship to GitHub by next weekend
Build MCP server for one daily-use tool Ship to GitHub by next weekend
Read 1 paper / week Public reaction on X or LinkedIn
Watch Anthropic + OpenAI + Sierra + Cognition job pages Weekly Friday review
Apply to OpenAI Residency 2026 Submit this month
Apply to Anthropic AI Safety Fellowship Submit this month
Apply to Google DeepMind Early Career Submit this month
Send 3 cold emails to frontier-lab engineers This week
Post 3-prediction Aluminium OS take on LinkedIn/X By Friday May 15
Ship MCP server + Claude Skill weekend project This weekend (May 17–18) — see 2026-05-13/03
Apply to 1 Anthropic FDE / Integration Engineer role This week — pair with the MCP artifact link
Read arXiv 2602.16666 + post 1 tweet/LinkedIn takeaway This week
Benchmark voice latency across Grok / OpenAI Realtime-2 / Google Speech By end of May
Make the focusing decision: commit to the Anthropic agentic stack This week — stop dabbling across ecosystems
Ship one OpenClaw skill + open the PR This weekend (4–6 hr project — see 2026-05-14/03)
Audit own model/token spend for 2 weeks Doubles as validation for the model-router startup wedge
Re-title resume in AI-native framing This week — same projects, +60–120% vs −40% conversion lane
Apply to 1 AI-infrastructure role This week — new less-crowded lane, Cisco data confirms demand
Watch Google I/O 2026 keynote (May 19) Googlebook/Aluminium OS SDK drop
Log where Meta May 20 layoff alumni land Start May 21 — leading signal for strong new product orgs
Pick ONE of the 5 AI sub-roles (Applied / Platform / LLM / Product / Responsible) This weekend — and rewrite resume headline to match
Ship + publish public MCP server (4–6 hr weekend project) By Sun May 17 night — pin above resume projects
Watch I/O keynote (May 19) and publish next-morning 1-page Gemini-vs-Claude-vs-OpenAI agent comparison By May 20 — single highest-leverage 4-hour artifact this week
Read "Attractor Models" + "Many Faces of OPD" end-to-end This week — earn the right to drop the vocabulary in interviews
Audit Claude programmatic spend before June 15 metering change ⚪ NEW — URGENT This weekend — see 2026-05-16/03
Retitle resume headline to "AI Integration Engineer in training" ⚪ NEW Tonight — see 2026-05-16/05
Apply to 5 specific AI Integration Engineer roles with MCP artifact attached ⚪ NEW This week — see 2026-05-16/05
Read "Cattle Trade" + "Successor-Representation Spectrum" papers ⚪ NEW This week — replaces last week's reading task once Attractor Models / OPD are done
Publish Gemini-vs-Claude-vs-OpenAI agent comparison Wed May 20 (post-I/O) Watch I/O Tuesday 10 AM PT; publish next morning
Pitch one local SMB on a "vertical-Claude-for-X" workflow library ⚪ NEW This week — doubles as customer discovery for the startup wedge
Drop CLAUDE.md (Karpathy template) into every active project root ⚪ NEW — TONIGHT Sunday May 17 — see 2026-05-17/03 §2
Enable prompt caching on highest-volume project (60–90% input savings) ⚪ NEW — TONIGHT Sunday May 17 — see 2026-05-17/03 §1; mitigation for June 15 metering change
Apply to 2 FDE roles tonight (Google Cloud + Anthropic) ⚪ NEW — TONIGHT Sunday May 17 — see 2026-05-17/05 §1; +800% FDE posting spike confirmed
Re-title LinkedIn headline to "AI Integration / FDE — Anthropic stack · MCP · Cost-aware agents" ⚪ NEW — TONIGHT Sunday May 17 — see 2026-05-17/05 §1; do before Monday-morning recruiter search wave
Pick ONE vertical-agent wedge and Loom-demo it to 1 buyer this week ⚪ NEW This week — see 2026-05-17/05 §3 — the rising-lane bet that pairs with the FDE fallback
Publish GRADED I/O comparison table (real Flash $1.50/1M numbers + own take) ⚪ NEW — TODAY Wed May 20 — see 2026-05-20/03 §1; grading your own prediction = credibility signal
Fix LinkedIn skills to real terms (Antigravity/Managed Agents/WebMCP), NOT "Vertex AI Agent Platform" ⚪ NEW — TODAY Wed May 20 — see 2026-05-20/01 §1
Ship dual-model "sanitiser" prompt-injection defense project (vulnerable vs defended traces) ⚪ NEW Fri May 22 — see 2026-05-20/05 §3; answers the #1 agent-deployment interview question
Ship WebMCP origin-trial demo (or "what I'll build when Chrome 149 lands" post) ⚪ NEW Sat May 23 — see 2026-05-20/03 §2; first-mover SEO repo of the month
Add Gemini 3.5 Flash as cheap leg in 3-provider router + per-step cost chart ⚪ NEW This week — see 2026-05-20/03 §4; now the most resume-relevant artifact
Add Google Cloud Agent / Antigravity Solutions roles to apply list (thin queue) ⚪ NEW This week — see 2026-05-20/05 §4
Pre-load Tuesday I/O viewing template (10 AM–1:30 PM PT) + 1-page Gemini vs Claude vs GPT post-keynote ⚪ NEW Monday/Tuesday May 18–19 — see 2026-05-17/03 §4
Toggle the Agent SDK credit setting in Claude account settings ⚪ NEW — TONIGHT Monday May 18 — see 2026-05-18/03 §2; credit doesn't auto-activate, silent fail June 15 if skipped
Pre-stage the Gemini vs Claude vs GPT comparison doc (pre-fill Claude Opus 4.7 + GPT-5.5 rows) ⚪ NEW — TONIGHT Monday May 18 — see 2026-05-18/03 §1
Pre-write LinkedIn / X teaser post template with 4 placeholders ⚪ NEW — TONIGHT Monday May 18 — see 2026-05-18/03 §1 Block 3
Block Tuesday 10 AM–5 PM PT for I/O viewing + comparison publish + 1 FDE application ⚪ NEW — TONIGHT Monday May 18 — see 2026-05-18/03 §1 Block 5
Set 12:45 AM PT alarm Wednesday for Code w/ Claude London Day-1 livestream ⚪ NEW Monday/Tuesday — see 2026-05-18/03 §3
Build "Thursday outreach short-list" of 10 Meta engineers (no contact yet) ⚪ NEW — TONIGHT Monday May 18 — see 2026-05-18/05 §1
Pre-write 5 Thursday DM templates (one per Meta sub-org archetype) ⚪ NEW Mon/Tue May 18–19 — see 2026-05-18/05 §1
Refresh LinkedIn headline to add "Vertex stack" ⚪ NEW — TONIGHT Monday May 18 — see 2026-05-18/05 §4 — signals readiness for Wednesday's Vertex Agent SDK hiring wave
Re-read ME.md + decide whether to add a vertical-PM-agent wedge to Active Portfolio ⚪ NEW Monday/Tuesday May 18–19 — see 2026-05-18/05 §5 and 2026-05-18/02 §4 — the most under-built of the "Suleyman 4" verticals
Apply to Isomorphic Labs engineering role (London / Cambridge MA / Lausanne) ⚪ NEW This week — see 2026-05-18/02 §1; rare AI-vertical employer not requiring biology PhD; lead with any comp-bio / chem-adjacent ML coursework
Read CHAL paper end-to-end + post 1-paragraph LinkedIn comparison to DyTopo + Bayes-consistent orchestration ⚪ NEW This week — see 2026-05-18/04 §1 — signals current-week multi-agent reading depth
Watch first 8 minutes of Sundar's I/O opener to classify consumer-first vs enterprise-first framing ⚪ NEW Tuesday — see 2026-05-18/01 §1 Insight — most predictive single signal for next-18-mo lab competition
Run 15-min-block I/O live-monitoring discipline + publish 1-page comparison by 12:30 PM PT ⚪ NEW — TODAY Tuesday — see 2026-05-19/03 §1 — single highest-leverage 4-hour artifact window of the month
Watch 12-min Code w/ Claude London slice (~1 PM PT) + publish 5 PM PT follow-up post ⚪ NEW — TODAY Tuesday — see 2026-05-19/03 §2 — asymmetric play of the day
Update LinkedIn headline + top-5 skills with post-keynote keyword by 11:55 AM PT ⚪ NEW — TODAY Tuesday — see 2026-05-19/05 §4 — 24-hour Wednesday recruiter-search window
Apply to one OpenAI FDE role before Tomoro-integration applicant flood ⚪ NEW Wed May 20 — see 2026-05-19/05 §2 — uniquely thin queue this week
Send 10 Meta-engineer DMs at 8 AM PT Thursday ⚪ NEW Thu May 21 — see 2026-05-19/05 §1 — outreach window opens with the May 20 layoff
Ship 1-evening AIRS-Bench portfolio project to GitHub ⚪ NEW Fri May 22 — see 2026-05-19/05 §3 + 2026-05-19/04 §1
Workday × Anthropic Solopreneurship Accelerator application ⚪ NEW Sat May 23 — see 2026-05-19/05 §5 — 15 slots, high signal
Build LinkedIn warm-channel to ~20 Tomoro FDEs (connect requests, no message yet) ⚪ NEW Wed May 20 — see 2026-05-19/05 §2
Build (mock) Tesco / Virgin Atlantic / Supercell FDE proposal 1-pager ⚪ NEW This week — see 2026-05-19/05 §2
Read JADE paper + ship a tiny claim-decomposition demo (2-day project) ⚪ NEW This week — see 2026-05-19/04 §2
Update STARTUPS.md wedge log with at least 3 new wedges this Saturday ⚪ NEW Sat May 23 — see STARTUPS.md
Start a personal apps/meta-alumni-tracker.md to log every Thursday DM + where they land 90 days out ⚪ NEW Today / Thursday — see 2026-05-19/05 §1