BK

Benjamin KassanStrategy & applied AI systems

Systems that
show their work.

I work on private-equity AI diligence and value creation at PwC Strategy&, build agent frameworks for consulting teams, and ship software where every number can be traced back to its source. Economics undergraduate; most of what I know I learned by shipping.

01Builds

Things I made, running here.

These are not screenshots. Each one runs the real logic on the real data from the project it came from — including the gaps, which are the most honest part.

Build 01

Radar Autonomy

A decision system for autonomous-fleet depot infrastructure — sourcing, diligence, underwriting, and a dossier that holds up when someone pulls on it. It is also the largest thing I have built: an agentic operating system with a product on top.

Set a fleet size and a budget. Twelve real San Francisco candidate sites get re-underwritten through the weighted gates. Hollow squares are unresolved inputs — they never quietly become a passing value.

Next.jsTypeScriptMapLibre GLGeospatial ETL
Advance0Conditional0Hold2Screen out10
3150 3rd St

PDR-2 · 51,597 sf · fits 128

87% evid.
6025 3rd St

PDR-2 · 94,304 sf · fits 235

87% evid.
2000 McKinnon Ave

PDR-2 · 242,498 sf · fits 606

100% evid.
180–200 Napoleon St

PDR-2 · 85,003 sf · fits 212

100% evid.
2025–2035 Newcomb Ave

PDR-2 · 32,000 sf · fits 80

100% evid.
2575 Marin St

PDR-2 · 27,000 sf · fits 67

100% evid.
2225 Jerrold Ave

PDR-2 · 98,614 sf · fits 246

100% evid.
450 Toland St

PDR-2 · 90,000 sf · fits 225

100% evid.
201 Toland St

PDR-2 · 64,000 sf · fits 160

100% evid.
1900 Oakdale Ave

PDR-1-B · 66,000 sf · fits 165

100% evid.
1313 Armstrong Ave

PDR-2 · 74,923 sf · fits 187

100% evid.
1300 Carroll Ave

PDR-2 · 121,097 sf · fits 302

71% evid.

Seven of nine production gates run here; site control and environmental need sources not shipped to this page, so the reachable weight is 79 of 100. Hollow squares are unresolved — they contribute no confidence and are never defaulted to a passing value. Hosting capacity is PG&E screening data and is not a service commitment.

Build 02

Prediction-market scout

Scores live prediction markets against a model-derived fair probability and only flags an edge when the divergence clears explicit thresholds. It ran across 407 markets and stored 424 assessments before the API budget ran out.

The diagonal is where the model agrees with the market. Everything interesting is the distance from it — and most of that distance is noise until you demand confidence too.

PythonClaudeSQLite
000.250.250.50.50.750.7511model agrees with marketMARKET PRICEMODEL FAIR PROBABILITY

Clears every gate

3 / 23

Hover a point. Its vertical distance from the diagonal is the divergence the screen is actually trading on.

23 real assessments from a corpus of 424 across 407 markets. Production defaults were |divergence| ≥ 0.06 and confidence ≥ 6. Single-model screen — the second-model consensus gate exists in the code but was never run.

Also built

01

Polymarket trading stack

Three generations of Python agents for prediction markets, roughly 17,000 lines. The engineering is the point: a separate risk engine and ledger, paper-versus-live parity tests, a websocket feed, a physical kill-switch file, and eleven test modules including evals for the LLM layer. It ran — 167,103 recorded market observations, 78 decisions, 59 executed trades across 33 markets.

PythonCLOB / websocketsSQLiteLLM evals
Private
02

Neuralese

A multi-agent trading research platform where four specialists — macro, quant, flow and risk — run on deliberately different cadences and write into one shared probabilistic world model. The risk agent can veto. It is the clearest thing I have built about agent coordination and structured disagreement, rather than about returns.

PythonMulti-agentFRED / market dataProvider-agnostic LLM adapter
Private
03

GlowAI

A finished agentic product loop rather than a chat window: baseline analysis, versioned plan generation, image projection through Gemini, weekly check-ins, and persistent memory that revises the plan and will not re-recommend something that already failed.

Next.jsTypeScriptGeminiFirebase
Private
04

In-browser AR face tracker

MediaPipe FaceLandmarker running entirely client-side at real-time frame rates, with three reticle styles, configurable EMA smoothing to kill landmark jitter, and live FPS and inference-time telemetry. Three dependencies, no backend, nothing uploaded.

TypeScriptMediaPipeViteComputer vision
Private
05

Newsglide

A news aggregation and summarisation service built on LLM workflows.

LLM workflowsWeb
06

American AI Dream

A site supporting American AI initiatives and education.

Next.jsClaude Code
02Hackathons

Build fast, then check honestly.

What I like about hackathons is the forcing function — and what I like about the one that went well is that the time went into building the thing that builds, rather than the thing itself.

Anthropic × Penn

Semi-finalist

November 2025 · Track 1, Agentic & MCP · with Adam Peles

Rather than build one project, we built the system that builds projects: five packaged Claude Skills and an orchestrator delegating to four domain builder sub-agents. It produced 36 complete projects — 411 files, 25,024 lines — and I wrote the rubric evaluator that scored twelve of them against the official judging criteria, plus the gallery the judges browsed them in.

Claude SkillsOrchestrationPythonFlask

Cook with Cursor

Build day

June 2026 · New York Tech Week · a build session, not a competition

Cursor’s NY Tech Week morning at The Malin. No tracks and no judging — you turn up and build. I spent it on a cross-venue arbitrage scanner for esports binary markets: pull active markets from two venues, normalise the claim types, fuzzy-match equivalents, read top of book, compute both hedge directions, and subtract fee and latency buffers before anything counts as a candidate. Read-only by default; every candidate needs a human to approve the settlement rules.

PythonPolymarket APIKalshi APIpytest

Top of the board: Neural Network Training Visualizer 9.67

01Neural Network Training Visualizer9.7
02Smart Code Review Agent9.7
03Momentum Trading Strategy Bot9.3
04Intelligent Document Q&A System9.3
05Personal Research Assistant Agent8.3
06RAG Chatbot with Memory System8.3
07CNN Image Classifier Visualizer8.3
08SQL Query Generator Agent8.0
09Semantic Search Engine7.7
10Mean Reversion Trading Bot7.7
11Attention Mechanism Visualizer7.7
12Portfolio Optimizer7.0
agenticragresearchtrading

Real scores from the rubric evaluator I wrote. The swarm built 36 projects (411 files, 25,024 lines); 12 were put through the evaluator, which is why 12 appear here and not 36. Change what the judge cares about and the winner changes — which is the honest thing a rubric does.

03Experience

Where I have worked.

Jun — Aug 2026

New York, NY

PwC Strategy&

Deal Technology Strategy

  • AI strategy across private-equity diligence and value-creation workstreams — assessing model, data, product and infrastructure maturity, translating technical risk into investment implications, and sequencing execution roadmaps.
  • Built reusable agent frameworks for consulting research and analysis, with source-of-truth hierarchies, specialist delegation, review gates and human approval boundaries.
  • Presented a practical agents playbook to 100+ Strategy& professionals on workspace design, multi-agent coordination and responsible adoption.

Sep 2025 — Present

Philadelphia, PA

Fingerpaint Group

AI Operations Intern

  • Built and deployed agents for paid-search optimisation and media-brief generation using Claude Code and Claude Skills, Gemini 3, and Antigravity.
  • Mapped end-to-end marketing workflows, identified the high-leverage automation points, and defined KPIs — time-to-brief, campaign performance, error rates — to hold the work accountable.

Sep 2025 — Present

Remote

Etienne Lane VC

Venture Capital Fellow

  • Built agent swarms that replicate the intern workflow: sourcing CPG and creator-led startups, scraping data, and synthesising one-pagers and investment briefs for partners.
  • Deployed sourcing and scheduling agents that identify promising founders, draft personalised outreach, and set up intro calls — widening top-of-funnel while freeing the intern class for deeper diligence.

Jun — Aug 2025

New York, NY

GradeWiz (YC-backed)

Student Success Engineer

  • Led the NYC private-school growth initiative: ran pilots in five schools end to end, documented results and teacher feedback, and signed two clients.
  • Shipped system improvements that cut latency and raised scoring agreement — prompt and rubric templates, OCR cleanup, LangChain batching and caching, and instrumented evaluation metrics built with the research team.

May — Aug 2025

New York, NY

Cornick, Garber & Sandler

Summer Analyst

  • Prepared individual and partnership filings, client write-ups and audit-prep packets for HNW clients; reconciled statements and standardised workpapers.
  • Mapped the return-prep workflow and piloted agentic automations — document intake to extraction to draft schedules — in n8n and React, with human-in-the-loop controls and audit trails, cutting manual entry and turnaround by up to 80% in practice.

Sep 2024 — Present

Philadelphia, PA

Education Consulting at Penn

Co-Founder & President

  • Founded and scaled a 40+ member consulting group working on education equity through data-driven strategy and grassroots outreach.
  • Led a nonprofit turnaround that raised membership 75%; ran an event series with SpaceX’s Ad Astra program drawing 60+ participants.
04Economics

What I actually think.

Mostly labour: what happens to employment when productivity moves, and why the answer is different in every industry. These are positions I hold and will defend, not a reading list.
01

Elasticity decides where the unemployment shows up

A productivity shock does not land evenly. If an industry faces inelastic demand, making its output cheaper does not expand the market much, so the gain arrives as the same output produced by fewer people. Where demand is elastic, cheaper output grows the market and employment can rise alongside productivity. So “will AI cause unemployment” is the wrong question — the elasticity is the parameter that decides, and it has a different value in every sector. Estimating it sector by sector is the actual work, and it is mostly not being done.

The through-line in how I read AI and labour

02

The studies measure a floor, not a ceiling

The credible causal work on AI and productivity is real and I accept it. But it studies suggestion-style tools from an earlier model generation, so the effect it identifies is a lower bound on what current systems do. Agentic systems move the question from “does this complement labour?” to “where does it substitute?” — and complement and substitute are not opinions about AI, they are empirical claims about a particular labour market.

My section of a labour economics presentation — “The Floor, Not the Ceiling”

03

Slower than I expected, and still a fast takeoff

I have been wrong on timing. Agents have arrived more slowly than I predicted, and I think the honest read is that the lag is in diffusion — org design, trust, procurement, who is allowed to approve what — rather than in capability. That distinction matters, because a capability overhang that adoption has not absorbed yet is not a reason to expect a gentle transition. I still think we are in a fast takeoff and that the economic change will be drastic; I have just stopped being confident about the quarter it shows up in.

A stated view, and a revised one

04

Grade your own evidence before someone else does

When I ran the analysis on whether AV-adjacent industrial property was actually repricing, most of my own results came back low confidence, my synthetic control was probably overfit, and one significant coefficient rested on too little data to interpret. I wrote all of that down next to the findings. An estimate without a confidence tier and a falsification condition is a decoration.

Self-directed difference-in-differences memo

Open invitation

I want to do actual empirical work on this.

If you are working on labour-demand elasticity, AI adoption and employment, or anything adjacent — as a researcher, a student, or someone sitting on data — I would like to hear from you. Co-authors, replication partners, and people who think I have this wrong are all equally welcome.

05Index

Everything else.

Education
University of Pennsylvania, B.A. Economics (2027). Minors in Statistics and Data Science. GPA 3.84. A year abroad at Queen Mary University of London reading labour economics and the political economy of food.
Leadership
Co-founder and President, Education Consulting at Penn — a 40+ member group working on education equity. Led a nonprofit turnaround that grew membership 75%, and ran an event series with SpaceX’s Ad Astra program.
Languages
Python · TypeScript / JavaScript · SQL · R · Mathematica
Systems
Claude Code & Agent SDK · agent orchestration · MCP · LangChain · n8n · Next.js / React · MapLibre GL · Postgres · Vercel
Methods
Difference-in-differences with staggered adoption · event studies · panel fixed effects · hedonic and count models · trees, random forests and gradient boosting · RNN forecasting
Away from work
National-level powerlifting · rugby · tennis · poker · chess · cooking

06Contact

Tell me what you are trying to decide.

I read everything. Diligence questions, agent-tooling problems, and anything about autonomy infrastructure get answered fastest.