- tl;dr sec
- Posts
- [tl;dr sec] #348 - Google's PageBreak Scanner, Perplexity's Agent Security Tool, Defending Agentically
[tl;dr sec] #348 - Google's PageBreak Scanner, Perplexity's Agent Security Tool, Defending Agentically
Google's AI-powered web app scanner with deterministic validation, OSS tool to detect, prevent, and investigate risky AI agent behavior, NCSC: "One does not simply defend agentically"
Hey there,
I hope you’ve been doing well!
💻️ Dev Day 2026
This week I attended my first ever OpenAI Dev Day, which was super cool!
It’s actually extremely difficult to go as an employee, very limited tickets, I just got very lucky (and shout-out to my friend Konstantine).
It was neat to feel the sparks in the air during the keynote, see the exhibit dedicated to builders, get a custom Dot sticker printed, check out the ImageGen photo booth, and see people building 3D cars in real time using Ultrafast mode.
One reflection I had during the event is that we spend a lot of time in security thinking about what can go wrong, and trying to prevent negative externalities. So it was refreshing to be surrounded by people who focus on the positive side of tech, and are just joyfully, playfully building.
It’s nice to be reminded of what we’re all working for.
And it was great to catch up with some creator friends like Jack Rhysider and Jasmine Wong, and make new friends. I also met Ben Burke at the Cerebras afterparty (from the KRAZAM comedic YouTube channel), whose Microservices sketch still delights me.
Quick funny story: I was chatting with some of the heads of security at OAI, and apparently earlier a PR person for a prominent YouTuber there came up to them and asked several times if they wanted a photo with the creator. They were like, “Nah I’m good,” and had no idea who the creator was. And the PR person had no idea who they were 😂 Delightful.
📣 Adopting Real-Time Threat Detection Workflows
Real-time threat detection is empowering security teams to flip the script on potential intruders. Real-time threat detection is the evolution of traditional threat detection that utilizes best-in-class modern security tools to analyze potential threats instantly. In the case of modern SIEM tools, teams can automate analysis to occur directly as event logs are ingested.
Learn how to adopt a real-time threat detection strategy with modern SIEM in under an hour.
👉 Learn More 👈
Congrats to my Panther friends for being acquired by Databricks 🙌 I’m excited to see the “security lakehouse” they build!
AppSec
Agentic Hacks, Real Proofs: Inside Google's PageBreak Project
Google's Michał Bentkowski writes about PageBreak, the Product Security team's agent, which has found over 500 XSS bugs across Google's first-party web applications since November 2025 (mostly using Gemini 3.1 Pro and 3.5 Flash), with near zero false positives. The key is their decision to prioritize deterministic validation: every hypothesis (potential vulnerability) goes to a validator that fires a real payload at a running environment, injecting JavaScript and watching whether it executes for XSS, writing a file in a world-readable location for path traversal, or triggering outbound DNS for RCE. Findings that cannot be deterministically confirmed act as seeds for deeper inspection in subsequent scans, and guide the development of new validation tools.
PageBreak also doubled as a test of the high-assurance web framework Google described last year, which is built to eliminate whole classes of web bugs by default. Across hundreds of applications built on it, the agent turned up 2 only XSS bugs, both in internal apps or debug endpoints with hardening gaps.
Michał credits Google's engineering setup for PageBreak’s effectiveness: the mono-repo lets the agent trace execution paths end to end without leaving the codebase, live HTTP traffic maps those paths back to source lines, and an existing scanner that authenticates to nearly every Google application lets PageBreak test many internal sites.
Google's PageBreak Project – Real-World Findings
Google's Michał Bentkowski walks through three examples of high-severity vulnerabilities PageBreak found in production Google applications, including a cache poisoning vulnerability, an XSS in admin.google.com, and a UXSS in the Tag Assistant Extension.
Sponsor
📣 The AI Playbook Security
Teams Need Right Now 📣
Employees are already using AI everywhere: in the browser, through extensions, inside vibe-coded apps, IDEs, MCP integrations, and agents running with their own permissions. Most security teams can't see any of it. This playbook breaks the AI attack surface into distinct layers, from browser-based AI to network-level AI traffic, and gives you pointed questions to test your visibility and control at each one. Use it to find the blind spots before attackers, auditors, or your CISO do.
Nice, I do like me some playbooks and checklists so I can quickly ramp up on what I should be keeping in mind.
Cloud Security
Mapping Scaleway IAM, An Attacker’s View of the Trust Boundaries
Tom De Keyser breaks down Scaleway's (an EU cloud provider) IAM model for anyone coming from AWS or Azure, where the concepts have the same names but don't work the same way. A Scaleway user principal API key doesn’t just carry permissions, it impersonates the person (leaked user key = full account takeover). Reconnaissance: even a policy-less key is enough to get oriented, leaking both metadata and permissions. For defenders, prefer application keys over user keys, guard the plaintext CLI config file, and assume all Organization and Project metadata is effectively public.
CiliumHound: Graphing Kubernetes Network Policies
SpecterOps's Andrew Luke introduces CiliumHound, a BloodHound OpenGraph extension for auditing Cilium network policies. (“Cilium is a Kubernetes Container Network Interface (CNI) plugin that enables network policies to control which pods can access resources within the cluster.”) CiliumHound ingests a folder of JSON or YAML policies and creates a searchable, Kubernetes namespace-scoped graph.
It addresses the challenge of auditing complex Cilium CNI policies by representing L3, L4, and L7 rules (different OSI model levels) as nodes and edges, with ingress rules flowing inward and egress rules flowing outward from central namespace nodes. CiliumHound includes saved Cypher queries to identify policy gaps like egress rules without L4 port restrictions.
Blue Team
Detection fidelity in the age of AI agents living off the land
Julien Vehent writes about why AI agents are breaking the tripwire detections (rare actions that legitimate users almost never take) defenders have relied on for years. Tools like curl, tcpdump, and throwaway Python scripts used to be rare enough that a single run was worth investigating, but agents now use them all day as part of normal work. He uses the Hugging Face incident as an example, where Hugging Face's sensors caught the activity but never paged anyone because each action looked like normal workload behavior. Taken together, the signals pointed to an intrusion, with credentials used from several environments, workloads minting tokens they'd never touched before, and more than 180 VPN enrollments.
Julien says detection has to move to correlation, giving each signal a confidence score and alerting when the combined score crosses a threshold along the attack chain. He also recommends giving every agent its own identity, testing that escalation actually reaches a human, and blocking agents from the cloud metadata service to cut noise at the source.
One does not simply defend agentically
The NCSC's Dave Chismon writes about why defenders can't adopt agentic AI the way attackers have, building on Halvar Flake's (Thomas Dullien) maxim that offensive problems are technical and defensive problems are political. Attackers chase clear success states like a working exploit, while defenders work through budgets, change approvals, and patch windows, with someone accountable for any automated action that takes the business down. Dave proposes scoring each defensive action on potency, scope, criticality, rollout confidence, and recoverability, to identify the lowest-risk actions that we can start to automate. Example easy wins include tasks that are low potency ('advise a human' rather than 'affect a system directly'), like summarizing threat intelligence.
The harder problem is proving an action really is low risk before it runs. “Can AI analyze traffic logs and show conclusively all the routes clients connect by? Can AI reverse engineer or otherwise assess a system and its binaries to show exactly which network calls it could ever make, or which processes it might need to spawn?”
Offense Scales with Compute. Defense Scales with Committees.
Bugcrowd founder and friend of the newsletter Casey Ellis writes about why AI is widening the gap between attackers and defenders. Attackers can now build and launch exploits as fast as their compute allows, while defenders still have to push every change through approvals and compliance first. Ellis says most AI-for-defense work focuses on the visible top of the stack, like AppSec and SIEM, which he calls "the top five turtles." Groups like the China-backed Volt Typhoon hide in the layers underneath, in router firmware, abandoned middleware, and old Windows XP machines nobody can update.
Shift from prevention to resilience by inventorying legacy systems, assume compromise wherever you lack visibility, and update incident response plans older than a year.
“The attacker's velocity of adoption is gated only by compute and creativity. The cost of failure is trying again. The defender's velocity of adoption is gated by enterprise policy, change control, production stability, compliance review, procurement, vendor consolidation, board appetite, insurance posture, and twelve people in four time zones agreeing on one Jira ticket. The cost of failure is loss of availability and revenue.”
💡 Excellent, thoughtful article. Warning: your wearable might record a slightly higher stress level after reading it.
AI + Security
Scan for Good: Using AI to discover and fix high-priority exposures across public services and critical infrastructure
Wiz's Ami Luttwak and Gal Nagli announce Scan for Good, which runs the Wiz Red Agent and Google DeepMind's Gemini 3.8 Flash Cyber to autonomously discover and help remediate exposures in public-facing systems of critical infrastructure, nonprofits, and public services. It appears they’ve disclosed ~475 High / Critical findings so far.
The post describes some high profile vulnerabilities found so far, including: an administrator key sitting on a public web server with read, write, and delete access to 8.8 million files in a national archive, a leaked production database exposed active admin sessions for a rail operator's routes and schedules, and a credential in public website code could have pushed malicious software into over 500 production container images behind a flagship AI service.
💡 Awesome that Wiz is putting some resourcing behind securing critical infrastructure, love to see it! 🙌
Securing Agents Across Perplexity’s Client Endpoints with Numbat
Perplexity open-sourced Numbat, an agent security suite that detects, prevents, and investigates risky AI agent behavior on macOS, Linux, and Windows. Numbat is a Go binary that plugs into harnesses like Claude Code, Codex, OpenCode, and Pi, using hooks to block risky actions before they run, session artifacts to rebuild timelines after the fact, and OpenTelemetry protocol (OTLP) telemetry for monitoring.
Those integrations feed 52 built-in rules, including sequence detections like flagging a secrets manager read followed by a curl upload, since either action alone can be legitimate. Rules use CEL expressions over normalized events. Inside Perplexity, Numbat runs on every endpoint through MDM and sends what those rules catch to Perplexity Computer, which investigates each detection and drafts new rules for the security team to review.
💡 Great post, and neat that they open sourced Numbat. Between that and Bumblebee, Perplexity has been cooking 🧑🍳 Shout-out Kyle Polley, Adel Karimi, and the rest of the team.
Misc
Misc
Anime Engineering Senpai - Software Engineer: I'm Tired of Pretending
Alex Hormozi - "Did I Waste My Best Years?"
Netflix trailer - One Year to Live, Buy a Man
Superhuman silky smooth movements - Pretty cool tricks
My First Million - Here's the (only) 8 rules for growing an audience
Humor
The Race For AGI - Crazy AI Mad Max parody with the foundation lab founders. I have mixed feelings about AI’s impact on screen industry workers, but the storytelling individual people can do now is impressive.
Ari K - CLAUDIO - A recruitment video
Kai Lentit - AI CEO Interviews (2026)
AI
As A.I. Makes Law Firms More Efficient, Clients Ask: ‘Where’s My Discount?’
Dwarkesh Patel - AI researchers debate how close we are to recursive self-improvement
AI Engineer - Jev Creator: Why RLCD beats RLHF
Brian Chesky - People managers won’t survive in the age of AI
Jeffrey Katzenberg (Walt Disney CEO, DreamWorks co-founder) - The World is Changing: AI For Creativity - Incredible story, great essay. “Technology didn't diminish the craft, it expanded the canvas. It gave artists more room to create.” “Tools are never the point. The instruments change with every generation. What endures is taste and imagination. The magical ability to make an audience feel.”
“As the barriers and the costs come down, more films will get made, not fewer. Studios will get to take more risks. There will be more seats at the table, and very soon entirely new forms of storytelling. In the 1980s, animation was dismissed as a niche corner of the business. Today it is one of the most beloved and profitable forms of storytelling in the world.”
✉️ Wrapping Up
Have questions, comments, or feedback? Just reply directly, I’d love to hear from you.
If you find this newsletter useful and know other people who would too, I'd really appreciate if you'd forward it to them 🙏
Thanks for reading!
Cheers,
Clint
P.S. Feel free to connect with me on LinkedIn 👋