Benjamin KassanStrategy & applied AI systems
Systems that
show their work.
I work on private-equity AI diligence and value creation at PwC Strategy&, build agent frameworks for consulting teams, and ship software where every number can be traced back to its source. Economics undergraduate; most of what I know I learned by shipping.
Things I made, running here.
Build 01
Radar Autonomy
A decision system for autonomous-fleet depot infrastructure — sourcing, diligence, underwriting, and a dossier that holds up when someone pulls on it. It is also the largest thing I have built: an agentic operating system with a product on top.
Set a fleet size and a budget. Twelve real San Francisco candidate sites get re-underwritten through the weighted gates. Hollow squares are unresolved inputs — they never quietly become a passing value.
PDR-2 · 51,597 sf · fits 128
PDR-2 · 94,304 sf · fits 235
PDR-2 · 242,498 sf · fits 606
PDR-2 · 85,003 sf · fits 212
PDR-2 · 32,000 sf · fits 80
PDR-2 · 27,000 sf · fits 67
PDR-2 · 98,614 sf · fits 246
PDR-2 · 90,000 sf · fits 225
PDR-2 · 64,000 sf · fits 160
PDR-1-B · 66,000 sf · fits 165
PDR-2 · 74,923 sf · fits 187
PDR-2 · 121,097 sf · fits 302
Seven of nine production gates run here; site control and environmental need sources not shipped to this page, so the reachable weight is 79 of 100. Hollow squares are unresolved — they contribute no confidence and are never defaulted to a passing value. Hosting capacity is PG&E screening data and is not a service commitment.
Build 02
Prediction-market scout
Scores live prediction markets against a model-derived fair probability and only flags an edge when the divergence clears explicit thresholds. It ran across 407 markets and stored 424 assessments before the API budget ran out.
The diagonal is where the model agrees with the market. Everything interesting is the distance from it — and most of that distance is noise until you demand confidence too.
Clears every gate
3 / 23
Hover a point. Its vertical distance from the diagonal is the divergence the screen is actually trading on.
23 real assessments from a corpus of 424 across 407 markets. Production defaults were |divergence| ≥ 0.06 and confidence ≥ 6. Single-model screen — the second-model consensus gate exists in the code but was never run.
Also built
Polymarket trading stack
Three generations of Python agents for prediction markets, roughly 17,000 lines. The engineering is the point: a separate risk engine and ledger, paper-versus-live parity tests, a websocket feed, a physical kill-switch file, and eleven test modules including evals for the LLM layer. It ran — 167,103 recorded market observations, 78 decisions, 59 executed trades across 33 markets.
Neuralese
A multi-agent trading research platform where four specialists — macro, quant, flow and risk — run on deliberately different cadences and write into one shared probabilistic world model. The risk agent can veto. It is the clearest thing I have built about agent coordination and structured disagreement, rather than about returns.
GlowAI
A finished agentic product loop rather than a chat window: baseline analysis, versioned plan generation, image projection through Gemini, weekly check-ins, and persistent memory that revises the plan and will not re-recommend something that already failed.
In-browser AR face tracker
MediaPipe FaceLandmarker running entirely client-side at real-time frame rates, with three reticle styles, configurable EMA smoothing to kill landmark jitter, and live FPS and inference-time telemetry. Three dependencies, no backend, nothing uploaded.
Newsglide
A news aggregation and summarisation service built on LLM workflows.
American AI Dream
A site supporting American AI initiatives and education.
Build fast, then check honestly.
Anthropic × Penn
Semi-finalistNovember 2025 · Track 1, Agentic & MCP · with Adam Peles
Rather than build one project, we built the system that builds projects: five packaged Claude Skills and an orchestrator delegating to four domain builder sub-agents. It produced 36 complete projects — 411 files, 25,024 lines — and I wrote the rubric evaluator that scored twelve of them against the official judging criteria, plus the gallery the judges browsed them in.
Cook with Cursor
Build dayJune 2026 · New York Tech Week · a build session, not a competition
Cursor’s NY Tech Week morning at The Malin. No tracks and no judging — you turn up and build. I spent it on a cross-venue arbitrage scanner for esports binary markets: pull active markets from two venues, normalise the claim types, fuzzy-match equivalents, read top of book, compute both hedge directions, and subtract fee and latency buffers before anything counts as a candidate. Read-only by default; every candidate needs a human to approve the settlement rules.
Top of the board: Neural Network Training Visualizer 9.67
Real scores from the rubric evaluator I wrote. The swarm built 36 projects (411 files, 25,024 lines); 12 were put through the evaluator, which is why 12 appear here and not 36. Change what the judge cares about and the winner changes — which is the honest thing a rubric does.
Where I have worked.
Jun — Aug 2026
New York, NY
PwC Strategy&
Deal Technology Strategy
- AI strategy across private-equity diligence and value-creation workstreams — assessing model, data, product and infrastructure maturity, translating technical risk into investment implications, and sequencing execution roadmaps.
- Built reusable agent frameworks for consulting research and analysis, with source-of-truth hierarchies, specialist delegation, review gates and human approval boundaries.
- Presented a practical agents playbook to 100+ Strategy& professionals on workspace design, multi-agent coordination and responsible adoption.
Sep 2025 — Present
Philadelphia, PA
Fingerpaint Group
AI Operations Intern
- Built and deployed agents for paid-search optimisation and media-brief generation using Claude Code and Claude Skills, Gemini 3, and Antigravity.
- Mapped end-to-end marketing workflows, identified the high-leverage automation points, and defined KPIs — time-to-brief, campaign performance, error rates — to hold the work accountable.
Sep 2025 — Present
Remote
Etienne Lane VC
Venture Capital Fellow
- Built agent swarms that replicate the intern workflow: sourcing CPG and creator-led startups, scraping data, and synthesising one-pagers and investment briefs for partners.
- Deployed sourcing and scheduling agents that identify promising founders, draft personalised outreach, and set up intro calls — widening top-of-funnel while freeing the intern class for deeper diligence.
Jun — Aug 2025
New York, NY
GradeWiz (YC-backed)
Student Success Engineer
- Led the NYC private-school growth initiative: ran pilots in five schools end to end, documented results and teacher feedback, and signed two clients.
- Shipped system improvements that cut latency and raised scoring agreement — prompt and rubric templates, OCR cleanup, LangChain batching and caching, and instrumented evaluation metrics built with the research team.
May — Aug 2025
New York, NY
Cornick, Garber & Sandler
Summer Analyst
- Prepared individual and partnership filings, client write-ups and audit-prep packets for HNW clients; reconciled statements and standardised workpapers.
- Mapped the return-prep workflow and piloted agentic automations — document intake to extraction to draft schedules — in n8n and React, with human-in-the-loop controls and audit trails, cutting manual entry and turnaround by up to 80% in practice.
Sep 2024 — Present
Philadelphia, PA
Education Consulting at Penn
Co-Founder & President
- Founded and scaled a 40+ member consulting group working on education equity through data-driven strategy and grassroots outreach.
- Led a nonprofit turnaround that raised membership 75%; ran an event series with SpaceX’s Ad Astra program drawing 60+ participants.
What I actually think.
Elasticity decides where the unemployment shows up
A productivity shock does not land evenly. If an industry faces inelastic demand, making its output cheaper does not expand the market much, so the gain arrives as the same output produced by fewer people. Where demand is elastic, cheaper output grows the market and employment can rise alongside productivity. So “will AI cause unemployment” is the wrong question — the elasticity is the parameter that decides, and it has a different value in every sector. Estimating it sector by sector is the actual work, and it is mostly not being done.
The through-line in how I read AI and labour
The studies measure a floor, not a ceiling
The credible causal work on AI and productivity is real and I accept it. But it studies suggestion-style tools from an earlier model generation, so the effect it identifies is a lower bound on what current systems do. Agentic systems move the question from “does this complement labour?” to “where does it substitute?” — and complement and substitute are not opinions about AI, they are empirical claims about a particular labour market.
My section of a labour economics presentation — “The Floor, Not the Ceiling”
Slower than I expected, and still a fast takeoff
I have been wrong on timing. Agents have arrived more slowly than I predicted, and I think the honest read is that the lag is in diffusion — org design, trust, procurement, who is allowed to approve what — rather than in capability. That distinction matters, because a capability overhang that adoption has not absorbed yet is not a reason to expect a gentle transition. I still think we are in a fast takeoff and that the economic change will be drastic; I have just stopped being confident about the quarter it shows up in.
A stated view, and a revised one
Grade your own evidence before someone else does
When I ran the analysis on whether AV-adjacent industrial property was actually repricing, most of my own results came back low confidence, my synthetic control was probably overfit, and one significant coefficient rested on too little data to interpret. I wrote all of that down next to the findings. An estimate without a confidence tier and a falsification condition is a decoration.
Self-directed difference-in-differences memo
Open invitation
I want to do actual empirical work on this.
If you are working on labour-demand elasticity, AI adoption and employment, or anything adjacent — as a researcher, a student, or someone sitting on data — I would like to hear from you. Co-authors, replication partners, and people who think I have this wrong are all equally welcome.
Everything else.
- Education
- University of Pennsylvania, B.A. Economics (2027). Minors in Statistics and Data Science. GPA 3.84. A year abroad at Queen Mary University of London reading labour economics and the political economy of food.
- Leadership
- Co-founder and President, Education Consulting at Penn — a 40+ member group working on education equity. Led a nonprofit turnaround that grew membership 75%, and ran an event series with SpaceX’s Ad Astra program.
- Languages
- Python · TypeScript / JavaScript · SQL · R · Mathematica
- Systems
- Claude Code & Agent SDK · agent orchestration · MCP · LangChain · n8n · Next.js / React · MapLibre GL · Postgres · Vercel
- Methods
- Difference-in-differences with staggered adoption · event studies · panel fixed effects · hedonic and count models · trees, random forests and gradient boosting · RNN forecasting
- Away from work
- National-level powerlifting · rugby · tennis · poker · chess · cooking
06Contact
Tell me what you are trying to decide.
I read everything. Diligence questions, agent-tooling problems, and anything about autonomy infrastructure get answered fastest.