Leonardo Palamim CardozoSoftware & AI Engineer
Summary

Six years in software engineering at early-stage startups, mainly on production TypeScript (React, Next.js, Node), with two businesses of my own along the way. I now work at the intersection of agentic coding and evaluation. I built an agentic coding tool by hand to understand agents below the frameworks, then built a lab to evaluate them, and published both the study and the reasoning that made me stop. Currently working on a floating persistent terminal and running an independent audit of the session_success metric in the SWE-chat dataset. Open to roles building or evaluating agent systems.

Recent Work

Starboard⁠

A terminal that's always beside your Dock. Not a Quake-style hotkey overlay, a permanent fixture.

A terminal that's always there. Starboard lives permanently beside the macOS Dock: on screen, on every desktop. There's no hotkey to call it up and no window to find, resize, or alt-tab back to, because it never goes away in the first place.

08/2026

swe-chat-audit⁠

An independent audit of the reliability of the session_success metric published in the SWE-chat dataset (Baumann et al., 2026, arXiv:2604.20779).

Findings and updates on my work⁠ page.

07/2026

magent-lab⁠

An offline evaluation and regression-testing suite for LLM-as-judge systems.

Versioned criteria and prompts, replicate runs, agreement measured against human labels, statistics implemented from scratch.

06/2026 – 07/2026

magent⁠

A planner/executor architecture with its own tool-calling loop, session orchestration, git integration, and a UI built for reviewing agent output before it ships.

Built by hand to understand agents below the frameworks. Used as my primary development tool. Retired it after evaluating it against frontier tools.

06/2026 – 07/2026
Publications

Why I stopped the conventions LLM-judge the study below had just validated. Code conventions resist decomposition into mutually exclusive criteria, and generating them for real collaborative repos (~US$16 for under 10% coverage across 4 trending repos) makes the approach unviable next to frontier tools.

07/30/2026

Independent, self-published. Measures whether an LLM-judge agrees with human labels and is self-consistent across 400 judgments (20 hand-labeled code diffs × 4 criteria × 5 replicates). Validity via majority-vote accuracy, sensitivity and specificity with Clopper-Pearson exact intervals; consistency via a split histogram of unanimity and Fleiss' kappa. Published with its limitations.

07/28/2026
Professional Experience

AI Builder & Researcher

Independent⁠

Full-time on agentic coding and evaluation. Projects, study and writing listed in Recent Work above; code on GitHub⁠.

05/2026 – Present

Founder & CEO

Forja⁠

Founded and ran a mentoring business for digital entrepreneurs. Grew to 8 mentees. Shut it down to return to engineering and AI.

06/2025 – 05/2026

Software Engineer

Zeca AI⁠

Shipped a retailer tool adopted daily by 100+ sales reps (DAU/WAU ~80%). Built features end-to-end across back-end, front-end, API integration and deployment. Stack: Node.js, TypeScript, Elasticsearch, React.js, Google Cloud, prompt engineering.

10/2023 – 06/2025

IR & Web Developer

Apex Capital⁠

Hired in Investor Relations; built and shipped the company's website⁠, and also a partner company's website⁠ (Next.js, TypeScript, Strapi, GraphQL, Docker).

10/2022 – 10/2023

Founder

Arbos Food

Founded an urban farming startup; raised R$160K from the Taqtile⁠ ecosystem. Didn't find product-market fit and folded.

08/2021 – 10/2022

Web Developer & Product Owner

Instituto Taqtile

Started as developer-in-training, grew into full-stack developer and product owner across client web and mobile projects. Next.js, React Native, Strapi, GraphQL, AWS.

11/2020 – 08/2021
Education

BSc. Completed alongside full-time roles and founding Arbos.

01/2017 – 12/2024
Skills
AI Engineering & Evaluation

Agents built from scratch; agentic coding (Claude Code); LLM-as-judge design; evaluation methodology; prompt engineering

Data & Infrastructure

PostgreSQL, Elasticsearch, Docker, AWS, Netlify

Languages & Frameworks

TypeScript, Python, React, Next.js, Node.js, Express, Prisma, GraphQL

Languages
Portuguese

Native

English

Fluent