GPT-5.6 Sol sets a new state of the art in cybersecurity on “The Last Ones” cyber range. We’re already seeing that capability translate into defensive outcomes: helping teams find, validate, and fix vulnerabilities in r…
Lesson: Defensive cyber SOTA ships as a product plugin (Codex Security) — evals that become workflows win.
GPT-5.6 Sol sets a new state of the art in cybersecurity on “The Last Ones” cyber range.
We’re already seeing that capability translate into defensive outcomes: helping teams find, validate, and fix vulnerabilities in real-world code.
Put it to work with Codex Security:
For the first time, China has taken the lead over the US in Frontend Code Arena with the launch of Kimi-K3 by @Kimi_Moonshot. The last time a Chinese model came close was in early 2025, with DeepSeek-R1.
Lesson: Arena leaderboards move markets: open Chinese models now lead frontend coding evals.
For the first time, China has taken the lead over the US in Frontend Code Arena with the launch of Kimi-K3 by @Kimi_Moonshot.
The last time a Chinese model came close was in early 2025, with DeepSeek-R1.
NEWS: SpaceX is reportedly in talks with the Pentagon to provide billions of dollars worth of data-center compute for its AI push. This deal could give the U.S. military access to SpaceX infrastructure, per Reuters.
Lesson: Defense AI compute is consolidating onto hyperscale + Starlink-adjacent infra players.
NEWS: SpaceX is reportedly in talks with the Pentagon to provide billions of dollars worth of data-center compute for its AI push. This deal could give the U.S. military access to SpaceX infrastructure, per Reuters.
NEW favorite artifact. I read this every morning to catch up on AI news from high-signal X accounts. It's an HTML artifact that curates X posts using the X MCP tools. Composed by my research agents. I have a daily aut…
Lesson: Personal morning digests from curated accounts + X MCP is the For You feed people actually want.
NEW favorite artifact.
I read this every morning to catch up on AI news from high-signal X accounts.
It's an HTML artifact that curates X posts using the X MCP tools. Composed by my research agents.
I have a daily automation that goes through curated X accounts and captures AI papers, projects, and more.
COG builds a self-evolving second brain using AI agents, Obsidian, and Git. It replaces the database with markdown files that sync via version control. https://github.com/huytieu/COG-second-brain
Lesson: Agent memory as git-backed markdown > opaque DBs for auditability and personal knowledge OS.
COG builds a self-evolving second brain using AI agents, Obsidian, and Git. It replaces the database with markdown files that sync via version control.
https://github.com/huytieu/COG-second-brain
Safe adopt prompt · repo / library / tool
Copy into Claude / Grok / Codex / Cursor — investigates provenance & malware first, then plans LifeOS-compatible install only if safe.
I tested two AI models on the same malware reversing challenge: the Temu version of Claude (Kimi K3) vs Claude itself. The challenge: analyze a Remus stealer using AI reverser skills - find the encrypted C2s and decrypt …
Lesson: Malware reverse evals are a real yardstick: fewer bigger tool calls can beat speed-only runs.
I tested two AI models on the same malware reversing challenge: the Temu version of Claude (Kimi K3) vs Claude itself.
The challenge: analyze a Remus stealer using AI reverser skills - find the encrypted C2s and decrypt them.
Claude Sonnet 4.6 (max) → 17:00 - 85 tool calls
Kimi K3 (max) → 25:00 - 37 tool calls
Both found obfuscated string tables, ChaCha20, and decrypted C2s. Kimi also found a fallback C2 behind an Ethereum contract (EtherHiding).
‼️ Microsoft wants to compete against Anthropic's Mythos with their 'Project Perception,' an AI security product that would route vulnerability-discovery and remediation tasks across Anthropic, OpenAI and Microsoft model…
Lesson: Multi-model vuln routing will become the enterprise default vs single-lab closed security products.
‼️ Microsoft wants to compete against Anthropic's Mythos with their 'Project Perception,' an AI security product that would route vulnerability-discovery and remediation tasks across Anthropic, OpenAI and Microsoft models.
They hope to give enterprises a lower-cost alternative to Anthropic’s restricted-access Mythos system.
AI can generate infrastructure. The harder problem is ensuring it follows your organization's standards. The Terraform MCP Server provides AI agents with authoritative context from Terraform workflows, modules, policies…
Lesson: IaC MCP servers are governance rails: agents only propose from approved modules + policy engines.
AI can generate infrastructure. The harder problem is ensuring it follows your organization's standards.
The Terraform MCP Server provides AI agents with authoritative context from Terraform workflows, modules, policies, and infrastructure configurations…
• AI-guided no-code infrastructure consumption
• Self-service deployments using approved registry modules
• Policy-aware workflows with Sentinel and OPA
• Enterprise-scale orchestration with Terraform Stacks
Lots of new stuff in recent Pipecat releases… three biggest new things: 1. Subagents 2. Folding Pipecat Flows into core 3. A new behavior evals framework. The LLM is the new function. The inference loop is the new threa…
Lesson: Production voice agents need subagent threads + explicit state machines (Flows), not one fat loop.
Lots of new stuff in recent Pipecat releases… three biggest new things: 1. Subagents 2. Folding Pipecat Flows into core 3. A new behavior evals framework.
The LLM is the new function. The inference loop is the new thread.
TaskManager and Workers support starting, stopping, and managing subagent threads with a shared message bus.
Ever wanted a powerful AI that runs right on your device? Meet Ternary-Bonsai-27B, a 2-bit compressed Qwen3.5 model that fits in your pocket, literally. It's designed for on-device chat and text generation. No cloud need…
Lesson: Aggressive quantization (2-bit) is making serious on-device privacy models practical.
Ever wanted a powerful AI that runs right on your device? Meet Ternary-Bonsai-27B, a 2-bit compressed Qwen3.5 model that fits in your pocket, literally. It's designed for on-device chat and text generation. No cloud needed, total privacy.
Spent my Friday out on the farm with @auterion. They’re shipping over 100,000 of their Skynode swarming & autonomous strike kits to Ukraine this year, including 50,000 with Ukrainian drone-maker Skyfall under a new deal…
Lesson: Autonomy kits at 100k+ scale: swarming autonomy is industrializing in Ukraine supply chains.
Spent my Friday out on the farm with @auterion.
They’re shipping over 100,000 of their Skynode swarming & autonomous strike kits to Ukraine this year, including 50,000 with Ukrainian drone-maker Skyfall under a new deal announced this week.
Lessons learned
Defensive cyber SOTA ships as a product plugin (Codex Security) — evals that become workflows win.
Arena leaderboards move markets: open Chinese models now lead frontend coding evals.
Defense AI compute is consolidating onto hyperscale + Starlink-adjacent infra players.
Personal morning digests from curated accounts + X MCP is the For You feed people actually want.
Agent memory as git-backed markdown > opaque DBs for auditability and personal knowledge OS.
Malware reverse evals are a real yardstick: fewer bigger tool calls can beat speed-only runs.
Multi-model vuln routing will become the enterprise default vs single-lab closed security products.
IaC MCP servers are governance rails: agents only propose from approved modules + policy engines.
Production voice agents need subagent threads + explicit state machines (Flows), not one fat loop.
Aggressive quantization (2-bit) is making serious on-device privacy models practical.
Autonomy kits at 100k+ scale: swarming autonomy is industrializing in Ukraine supply chains.