Andrew Borowicz

I build the software my business runs on.

I'm a finance student at Trinity University, graduating May 2027. I run Lone Pine, an Amazon resale business that did over $1M in sales in the last 12 months on an operations platform I built, and I co-founded Mass Apply AI, a desktop app in private alpha that fills in job applications for you. Claude writes most of the code; I decide what gets built and check it against the numbers.

$1M+Lone Pine sales, last 12 months
~$150knet profit, after overhead
25Mass Apply releases since May
5×more roles per sweep after one fix

Selected work

Every project below runs live on this page

Lone Pine

Jun 2023 – now

My Amazon business, and the operations platform I built to run it.

The real app with sample data. Flip through the screens, or choose Try it live to open orders, trace a unit and build a claim packet.

Lone Pine
Lone Pine Briefing screen with sample data. The start-of-day page: what needs attention, ranked, with the week in five numbers underneath.

The start-of-day page: what needs attention, ranked, with the week in five numbers underneath.Real app · sample data · scroll sideways

The business trailing 12 months

$1M+in Amazon sales
~$150knet profit
~50%gross margin, ~15% net
30,000+units sold
$50k$100kOctNovDecJanFebMarAprMayJunJulAugSep

Sep 2026 (through the 27th) · ~$113k sales · ~$17.7k net profit

Sales by month, net profit as the lower bar · Sep still settlingHover or tap a monthTap a month

Figures come from the platform's own P&L, which counts real unit costs, Amazon fees and overhead. Rounded on purpose. Sales data in the platform starts February 2025; 3,500+ purchase orders from 200+ suppliers since then.

Lone Pine buys products from retailers and resells them on Amazon. It first ran on Airtable. I replaced that with a Next.js and Postgres app covering purchasing, inventory, reimbursements and accounting, and moved all 2,671 purchase orders on file at the time, with nothing missing.

The rule behind every screen: show the real number, or show nothing. Every unit gets a code when it's bought and is tracked until it sells. Every card charge is matched to a purchase order, or the app says in plain English why it couldn't be. Claim packets that took about 110 minutes a week to put together by hand now build in one click.

In September I redesigned the whole app to feel like a native Apple app. The brief came from real usage: in one month our sourcing VA created 324 purchase orders and pasted 122 tracking numbers, about 40 edits a day on one screen, and the team missed how instant Airtable felt. So it had to be at least that fast. Cold load on the main tab went from 23.2 s to 4.8 s, and first-load JavaScript fell by 60%.

Commits
1,637
Merged pull requests
183
Dashboard tabs
42
Database migrations
208
Per-unit ledger
Live since Aug 1
Scheduled jobs
44
First-load JavaScript
−60%
Stack
Next.js · Postgres · Vercel

The demo uses sample data. Suppliers and staff stay private.

How it works Next.js · Postgres · Vercel

SourcesAmazon's seller API, Gmail, Plaid card feed, shipping carriers, the prep center and a repricer.
44 scheduled jobsSync every source on a schedule, each one watched so a dead feed can't look healthy.
Postgres208 migrations, with row-level security so each person sees only what their role allows.
EnginesPer-unit ledger, card-charge matching, claim packets and a ranked daily task queue.
42 tabs + SlackThe team works in the app; approvals and digests arrive in Slack.

Four hard problems tap to open

Profit per unit, when Amazon never says which unit soldOne 34-second job
Problem
Amazon reports that a product sold, not which of the units you bought it was. Without that, profit per unit is a guess.
Fix
Every purchased unit gets a code with its real cost frozen at purchase. Sales use up the units that evidence shows reached Amazon first, so a lost unit can never be mistaken for a sale on paper.
Result
The first run traced 36,923 units (bought, in stock or sold) in a single 34-second job. Sales it can't attribute are counted and shown, never guessed.
Lesson
Comparing timestamps as text would have rewritten every row every night. They're compared as real times now.
Claim packets were missing invoices~110 min a week automated
Symptom
Reimbursement claims came out without the supplier invoices that prove them.
Cause
The app searched a synced copy of the inbox, which only held a capped subset of emails.
Fix
Search Gmail live at the moment the packet is built, render each email exactly as it looks, and fill the claim letter automatically.
Result
Replaced about 110 minutes a week of manual work for our reimbursements VA. A second review caught two ways the tracking-number reader could have made one up.
A dashboard that looked green while brokenFailures now show as failures
Approach
A tab-by-tab review, with every finding re-measured before anything was built on it.
Found
Sections showed "all clear" when their data had actually failed to load, and a hidden 1,000-row limit was producing wrong overdue counts.
Result
Every broken section now says so. Cold load for the main tab dropped from 23.2 s to 4.8 s, mostly by deleting duplicate reads.
Lesson
One proposed alert threshold came from a measurement mistake. Re-measuring moved it from 16 days to 8.
Making it fast60% less JavaScript
Change
Cut first-load JavaScript from 495 KB to 193 KB and moved the servers next to the database.
Not shipped yet
Rewriting 112 security rules took one slow query from 1,037 ms to 21 ms in testing. It is ready and waiting on my final review, so it isn't counted above.

Timeline

  • May 27First commit and first deploy
  • Jun 10Airtable retired: 2,671 orders moved, nothing missing
  • Jun 16One-click claim packets
  • Aug 1Per-unit ledger goes live
  • Aug 3Live card feed
  • Sep 22Full redesign in a native Apple style
  • Sep 27183rd merged pull request

What I'd do differently

  • Build a health signal for every data feed on day one. One feed was dead for weeks while reporting green.
  • Walk through every staff workflow after a redesign. Six buttons were white-on-white for five days before anyone told me.

I own the business and wrote the platform. Our two VAs, one on sourcing and one on reimbursements, shaped it with daily feedback.

Mass Apply AI

May 2026 – now

Co-founded the company and built the desktop app, which ships as Minsky. Private alpha since Sep 6.

The real app with sample data. Flip through the screens, or choose Try it live to run an application and answer the question it parks for you.

Mass Apply AI
Job Feed: 16,000+ roles ranked by match score, with filters, sponsorship and an Auto-Apply button per row

Every opening from a sweep, ranked against your profile. Auto-Apply where there is an adapter, Fill manually where there isn’t.Real app · sample data · scroll sideways

Mass Apply AI finds openings across thousands of company career sites, ranks each one against your profile, and fills in the application. It never presses Submit. When a form asks something only you can answer, like relocation or salary, it leaves the field blank, asks you once, and reuses the answer on every form after that.

Coverage was the hardest part: a hidden bottleneck skipped about 1,300 job boards on a typical run, and fixing it made each sweep return five times as many roles. The full story is in Notes.

Roles per sweep
3,250 → 16,310 (Sep 17)
Boards reached
99% of 2,971
Workday forms reaching Review
1 → 13 of 20
Application platforms
11
Releases shipped
25
Tests
4,587
Commits
442

The screens are the shipping app rendered in a browser with sample data: fictional employers and a fictional applicant. “Try it live” uses public postings from a real feed capture, simplified.

How it works Electron · React · TypeScript · SQLite · Playwright

FindPulls openings from 2,971 job boards across the big hiring platforms, and adds new boards it discovers on its own.
StoreLocal-first: jobs, your profile and every answer live in a SQLite database on your Mac.
RankA score across skills, level, sector, location, sponsorship and freshness that learns from what you apply to and skip.
FillEleven adapters, one per platform, drive a real browser. Answers come from your profile first; AI only drafts open-ended text.
StopA guard blocks every Submit click and Enter press. The form waits at Review until you submit it.

Four hard problems tap to open

1,300 boards a run were silently skipped5× more roles per sweep
Symptom
About 1,300 boards on a typical run were never even tried, and nothing reported an error.
Cause
A circuit breaker counted failures per web address. Every Greenhouse, Lever and Ashby board shares one, so a few slow employers switched off a whole platform.
Fix
A higher threshold, a short retry probe, and a meter that counts what a user would actually see.
Result
Five times as many roles per sweep; 99% of 2,971 boards reached.
Workday forms stalled before Review1 → 13 of 20 employers
Symptom
Workday runs a different form on every employer’s site, and most runs died before the Review step.
Fix
Typeahead pickers that type, press Enter and keep Workday’s own pick; sign-in handling; a test harness across 20 real employers.
Result
Employers reaching Review went from 1 of 20 to 13 of 20.
Constraint
One full live test a day. Repeated sign-ins lock the test account, so every run had to count.
A wrong answer is worse than a blankRelease held until fixed
Approach
Ran the answer rules over live application forms, filling but never submitting, and read what they chose.
Found
A backwards sponsorship answer, school names landing in email fields, and a “Not Hispanic” option matching “Hispanic”.
Fix
Held the release until each one was fixed. When the app isn’t sure now, it leaves the field and asks you once.
Lesson
For anything that speaks for a person, careful beats complete.
A fix I shipped that broke thingsReverted and repaired
What happened
I changed Greenhouse apply links to a cleaner format, checked them with a command-line request, and shipped. On embedded job boards, that page quietly redirected to a site with no form.
Repair
Reverted the change, wrote a database migration to fix the links those releases had saved, and opened the next release notes with the correction.
Lesson
Read the current code before acting on an old bug report.

Timeline

  • May 31First commit
  • Sep 1Seven application platforms filling reliably
  • Sep 6Private alpha ships, with an in-app updater
  • Sep 8Workday: 13 of 20 employers reach Review
  • Sep 17Coverage fix: 5× more roles per sweep
  • Sep 2825th release

What I’d do differently

  • Verify every fix the way a user runs the app, in the real browser and the packaged build, before calling it done.
  • Script the release process on day one. An early release raced two uploads and shipped broken.
  • Send every platform through one shared answer pipeline from the start, instead of letting Workday keep its own.

Built with my co-founders Jasper Buntinx (web and funnel) and Diego Prozzi (growth). I wrote the desktop app. About 2,100 of the 2,971 job boards came from Jasper’s board list.

Résumé

Interactive · tap anything
Tap a role to see what I did there

Education Expected May 2027

Trinity University

Neidorff School of Business · Bachelor of Business, Economics and Finance · San Antonio, TX

3.58cumulative GPA
Dean's List2024–26
2027expected graduation

National Hispanic Recognition Program · AP Scholar

Finance and investing

Equity Valuation · Principles of Investing · Real Estate and Alternative Investments · Corporate Finance · FINRA Securities Industry Essentials

Accounting

Financial Accounting · Intermediate Accounting I · Intermediate Accounting II

Quantitative and business

Advanced Spreadsheet Modeling · Statistics for Business and Economics · Business Law · Business Management · Operations Management

Skills tap one to see where I used it

Organizations

San Antonio FPA · Chess Club · ModernGuild · Youth on Course Golf · Trinity Men's Soccer Club · Armadillo Bouldering

Taska

Aug – Sep 2026

Built my own command center: school, calendar and training in one ranked list.

The real app, running on its built-in demo courses. Choose Try it live to check items off and watch the lock-screen widget update.

Taska
Taska week plan: each day's work blocks, capped at three hours a day, with due dates

The week, planned automatically: every deadline split into blocks under a daily cap.Real app · demo courses · scroll sideways

Taska pulls in Canvas assignments, calendar events, GitHub and my training plan, turns deadlines into a week of work blocks under a daily cap, and serves the top of the list to a lock-screen widget on my phone. It's a single Node file with no dependencies, plus an MCP server so Claude can read and update my to-dos directly.

The first version could hang for over a minute waiting on one slow source. Now every source gets its own deadline and a ten-minute cache: a cold load takes 2 seconds and a warm one 12 ms. It started life as "Canvas Planner", which is still the name in its sidebar.

Screens use the app's demo mode, with sample courses. My real classes and accounts stay private.

Load, cold / warm
2.0 s / 12 ms
MCP protocol tests
20 / 20
Dependencies
0
Sources
Canvas · calendar · GitHub
Stack
Node · MCP · iOS widget

My first IRONMAN 70.3

Jul – Sep 2026

A training plan with every minute of every day accounted for.

The real site, one Saturday in the peak week. Choose Try it live and pick any block: each day has to add up to exactly 1,440 minutes or the build fails.

myfirstironman
Training site day view for a Saturday: date strip, a 24-hour ribbon and every block from pre-ride breakfast to sleep

One day, start to finish: a long ride, the run off the bike, meals, rest and sleep.Real site · personal details removed · scroll sideways

I set out to train for my first half-IRONMAN (a 1.2-mile swim, 56-mile bike and 13.1-mile run), aimed at Waco in October, on top of a full course load and a business. There's no slack in that week, so I treated the plan like software: one data file, and a Python build that fails unless every day covers exactly 1,440 minutes. It publishes a five-page site, a 141-event calendar feed and a 5:45 morning push to my phone.

The swim got dates instead of opinions: miss the 1,500-yard checkpoint and the plan defers to a spring race rather than force it. In the end I didn't register for Waco, so the plan rolls forward to a spring race, and the build recomputes every date.

Next target
A spring 2027 70.3
Minutes per day, checked
1,440
Calendar events
141
Swim checkpoints
3, dated
Pages
Daily · Training · Diet · Gym · Progress
Build
Python → static site

Screens are from the real site with nutrition targets and locations removed. “Try it live” is a sample week.

JARVIS

May – Sep 2026

A personal AI memory system, tested the way researchers test them.

Step through how a memory gets from my day into an answer.

JARVIS
How it works

Notes, decisions and people live as plain markdown files in one vault. No database to lose.

Scheduled jobs write a daily and weekly summary, so recent context is always condensed.

A router reads each question and picks how the retrieved memories are laid out for it, instead of one format for everything. It still misroutes most questions that span many chats (right on 5 of 25).

A custom MCP server lets Claude search the vault directly, from any chat.

LongMemEval · 150 questions
0.767115 of 150 correct
Baseline0.713
With my router0.767
By question type, 25 eachbaselinerouter
  • Assistant said, one chat1.00 → 0.96
  • User said, one chat0.80 → 0.88
  • Preferences0.24 → 0.44
  • Facts that changed0.72 → 0.84
  • Across many chats0.68 → 0.72
  • Dates and order0.84 → 0.76

A public benchmark for long-term memory in chat assistants. The full run cost $4.30 of a $5 budget. The router helped most on preferences and changed facts, and made date questions worse.

I wanted Claude to remember my work across weeks, not just within one chat. JARVIS is a second brain in plain markdown, summarized every day by scheduled jobs and searchable through an MCP server I wrote.

The interesting part was proving it worked. I ran it on LongMemEval, a public benchmark, added a router that picks a strategy per question, and measured the gain on a strict budget. I wrote down what would count as success before the run, and 0.767 landed in the middle band: the router works, just less than I'd hoped.

Then I looked at where the remaining misses come from, by comparing against a run that is handed exactly the right memories. About 62% of that gap is in reading the answer, not finding it, so better search alone won't close it.

Benchmark score
0.713 → 0.767
Running
Daily since May 9
Vault commits
194
Eval spend
$4.30 of $5

Also built

2026
Design · Lone PineA design system that feels native

Native type, comfortable density and one idea per screen, written as a spec and rolled out across every tab at once.

359 commits in one redesign
AI operationsMorning business recap

Every morning, three read-only agents go through Slack, the database and Amazon's API changes, and write one short, ranked recap.

3 agents · read-only · daily
Study toolsCourse review pipeline

One agent per lecture deck turns a module into a quiz-review site, and every number in it is recomputed in Python before I trust it.

7 lectures · 1 module
Quantitative researchMarket efficiency study

Priced NFL season-long player props against the market. The honest finding: it's nearly efficient.

1,318 props · 282 players
Finance · Lone PineLive card feed

Two years of card transactions pulled in through Plaid and matched to suppliers, each match with a written reason.

4,712 transactions
Code qualityMulti-agent audit

A tab-by-tab review of the Lone Pine app, with every finding re-measured before anything was built on it. The worst find: a data feed that had been dead for weeks while its job kept reporting success.

82 findings

Notes

Short write-ups on what I learned
Sep 2026The bug that skipped 1,300 job boards a runA safety feature quietly turned off whole hiring platforms, and nothing looked wrong.4 min

Mass Apply AI sweeps thousands of company career pages every day. The feed felt thin, but nothing was failing loudly. The logs said the sweep finished. The app showed jobs. Everything looked fine.

So I stopped reading the code and started counting. On one sweep, 1,303 of 1,491 board failures weren't failures at all. Those boards had never been tried.

The cause was a circuit breaker, the kind of safety switch that stops you from hammering a server that's struggling. Mine tracked failures per web address. But every Greenhouse board lives at the same address, and so does every Lever board and every Ashby board. Three slow employers were enough to switch off an entire platform for the rest of the run.

1,303boards never tried, on one sweep
5×roles per sweep after (16,000+)
99%of 2,971 boards reached

The fix was small: a higher threshold, a 30-second retry probe, and dropping about 1.3 GB of page data per sweep that the app didn’t need. The bigger change was a tool that measures what a user would actually see, so the next silent failure shows up as a number.

What I took from it: a job that runs once a day never shows you the part it keeps skipping. If you don't measure coverage directly, you are trusting the absence of errors, and that isn't the same thing as success.

Aug 2026A dash beats a guessWhy my business dashboard shows "—" instead of a number it isn't sure about.3 min

When I started building Lone Pine's platform, I wrote one rule at the top of the project: show the real number or show nothing. A confident wrong number is worse than a blank, because people act on it.

That rule shaped the hardest part of the system. Amazon tells you a product sold. It never tells you which unit sold. The easy answer is first-in, first-out. But then a unit lost in a warehouse gets quietly matched to a sale on paper, and your profit looks better than it is.

Instead, every unit gets a code with its real cost frozen when we buy it. A sale can only use up a unit that evidence shows actually reached Amazon. The first run traced 36,923 units in 34 seconds. Sales it can't attribute are counted and shown, not guessed.

The same rule applies to profit. When a product's cost was never entered, those sales are left out of profit instead of estimated. It means the dashboard slightly understates what we make. I'm fine with that. An understated number I can trust is worth more to me than a flattering one I have to double-check.

Sep 2026How I work with ClaudeFour habits that made AI-written code something I could actually ship.4 min

Claude writes most of my code. That makes my job less about typing and more about judgment: deciding what to build, and knowing when something is actually done. Four habits made the difference.

Measure what the user sees. Twice, reading the code convinced us that a style was broken. Twice, measuring it in a real browser proved otherwise. Now any claim about the interface gets measured before anyone acts on it.

Split the work. Independent pieces run in parallel on separate branches, then merge only once each one stands on its own. That's how seven page redesigns landed in a single day.

Look for what’s wrong on purpose. Before one release, I had a separate agent review the build with a single question: where could this app answer wrongly for someone? It found seven answers the app would have filled in incorrectly on real job applications. The release waited until all seven were fixed.

Own the misses. I once shipped a link "fix" I had only checked with a command-line request. In a real browser it broke embedded job boards. I reverted it, wrote a migration to repair the damage, and led the next release notes with the correction. The lesson is now a rule: check it the way the user will.

Contact

San Antonio, Texas

Internships and full-time roles in product or finance, or a problem worth solving. I read everything.

hello@andrewborowicz.dev