Six years in software engineering at early-stage startups, mainly on production TypeScript (React, Next.js, Node), with two businesses of my own along the way. I now work at the intersection of agentic coding and evaluation. I built an agentic coding tool by hand to understand agents below the frameworks, then built a lab to evaluate them, and published both the study and the reasoning that made me stop. Currently working on a floating persistent terminal and running an independent audit of the session_success metric in the SWE-chat dataset. Open to roles building or evaluating agent systems.
A terminal that's always there. Starboard lives permanently beside the macOS Dock: on screen, on every desktop. There's no hotkey to call it up and no window to find, resize, or alt-tab back to, because it never goes away in the first place.
Findings and updates on my work page.
Versioned criteria and prompts, replicate runs, agreement measured against human labels, statistics implemented from scratch.
Built by hand to understand agents below the frameworks. Used as my primary development tool. Retired it after evaluating it against frontier tools.
Why I stopped the conventions LLM-judge the study below had just validated. Code conventions resist decomposition into mutually exclusive criteria, and generating them for real collaborative repos (~US$16 for under 10% coverage across 4 trending repos) makes the approach unviable next to frontier tools.
Independent, self-published. Measures whether an LLM-judge agrees with human labels and is self-consistent across 400 judgments (20 hand-labeled code diffs × 4 criteria × 5 replicates). Validity via majority-vote accuracy, sensitivity and specificity with Clopper-Pearson exact intervals; consistency via a split histogram of unanimity and Fleiss' kappa. Published with its limitations.
AI Builder & Researcher
IndependentFull-time on agentic coding and evaluation. Projects, study and writing listed in Recent Work above; code on GitHub.
Founder & CEO
ForjaFounded and ran a mentoring business for digital entrepreneurs. Grew to 8 mentees. Shut it down to return to engineering and AI.
Software Engineer
Zeca AIShipped a retailer tool adopted daily by 100+ sales reps (DAU/WAU ~80%). Built features end-to-end across back-end, front-end, API integration and deployment. Stack: Node.js, TypeScript, Elasticsearch, React.js, Google Cloud, prompt engineering.
IR & Web Developer
Apex CapitalFounder
Arbos FoodFounded an urban farming startup; raised R$160K from the Taqtile ecosystem. Didn't find product-market fit and folded.
Web Developer & Product Owner
Instituto TaqtileStarted as developer-in-training, grew into full-stack developer and product owner across client web and mobile projects. Next.js, React Native, Strapi, GraphQL, AWS.
Mechanical Engineering
Polytechnic School of the University of São Paulo (Poli-USP)BSc. Completed alongside full-time roles and founding Arbos.
Agents built from scratch; agentic coding (Claude Code); LLM-as-judge design; evaluation methodology; prompt engineering
PostgreSQL, Elasticsearch, Docker, AWS, Netlify
TypeScript, Python, React, Next.js, Node.js, Express, Prisma, GraphQL
Native
Fluent