A career site monitor built almost entirely by AI agents, in seven days
Job Watcher checks a configurable list of company career sites, classifies new postings with a local AI model, and emails a report every morning by 7 AM. An orchestrating agent planned the build and directed roughly twenty five subordinate agents to write it, test it, and fix what broke.
Told what to build, not how to build it
The project started with three specification documents describing what the product had to do. Deliberately, they left out how. The core requirements were strict:
- Monitor a configurable list of company career sites. Companies could not be hard coded into the source.
- Identify newly discovered postings and never report the same job twice.
- Classify each new job as AI, Graphic Design, Both, or Neither, using a local Ollama model judging actual responsibilities rather than title keywords.
- Extract any stated security clearance requirement. Report it, never use it to filter results.
- Email a report every morning by 7:00 AM, including mornings with nothing to report.
- Run with no Terminal open and no application window open.
- Tell the difference between a source that failed and a source that genuinely has zero new jobs.
- Stay out of the applicant matching business entirely. No fit scores, no qualification judgments.
Eleven starter companies were specified, ranging from defense contractors to a grocery conglomerate to a federal research institute.
How It Works
Once configured, Job Watcher runs the same routine every day without anyone opening the app.
- Checks each configured career site through the platform adapter that matches how that company posts jobs.
- Compares what it finds against everything already stored, so a posting is only ever reported once.
- Sends the full text of each new posting to a local Ollama model, which classifies it as AI, Graphic Design, Both, or Neither and writes down its reasoning.
- Pulls out any stated security clearance requirement and stores it for reference. It never uses it to decide what gets reported.
- Builds the morning report from whatever is relevant and sends it by email before 7 AM, even when there is nothing new.
One Orchestrator, Many Agents
One Opus class agent ran the project as coordinator and almost never wrote code itself. It read and protected the specification, researched the problem space before choosing an architecture, and broke the work into pieces with strict file ownership so agents running in parallel could never overwrite each other. Implementation, research, testing, and debugging went to roughly twenty five Sonnet class subagents. The orchestrator reviewed and independently checked every result before accepting it, and only did the work directly when delegating it was not possible.
That file ownership discipline earned its keep once. A rate limit killed three agents in the middle of their work at the same time. Because each one owned a separate set of files, nothing was left in a corrupted state. The damage was five missing modules and one failing test, both easy to find and finish.
Midway through the build, the orchestration strategy was adjusted after running several agents concurrently triggered rate limit failures. Three or four agents running at once turned out to be the trigger, so the approach shifted to fewer agents working at a time, eventually favoring sequential delegation over parallel execution for reliability. One agent working at a time was slower per step, but it never lost work.
Agents were also told repeatedly to report a problem honestly rather than patch around it, especially when a fix belonged in another agent's files, or when the honest answer was that something could not be done.
This orchestrator-and-subagent structure, the reason it holds up under real failures, is discussed in more general terms in Why Most AI Agents Break. Job Watcher is that architecture applied to a real build.
Architecture
The build settled on a short list of deliberate, boring choices.
| Concern | Choice | Why |
|---|---|---|
| Language | Python 3.13, project local virtual environment | No runtime to install |
| Storage | One SQLite file, WAL mode | Zero infrastructure, safe for a UI and a scheduled run writing at the same time |
| Scheduling | launchd LaunchAgent | Survives a reboot, needs no Terminal or open window |
| AI | Local Ollama (gpt-oss:20b) over HTTP | No API cost, no data leaves the machine |
| Gmail API, gmail.send scope only | Cannot read the inbox | |
| UI | Local web app bound to 127.0.0.1, plus a menu bar agent | No Electron, no cloud |
No per company code. Companies are rows in a database, not source you have to touch.
Collection happens through generic platform adapters, with company specifics living in configuration data. That rule is what makes "add a company without touching source code" true instead of aspirational. Four additional companies were added through the interface during production testing, without modifying source code.
The Numbers
Figures below were checked against the repository and the live production database on September 22, 2026. None of them are estimates.
Research Before Code
Before writing any collection code, agents worked out how each company's career site actually behaves, checking every endpoint with live requests instead of trusting documentation.
A second research pass solved two sources that looked impossible. One defense contractor's job board returned an HTTP 500 error until an agent found a hidden request verification token that the site's own JavaScript forwards as a header. Using it returned 782 openings. A federal research institute's jobs turned out to be published on USAJOBS rather than the institute's own careers page, and a keyless endpoint was found there, removing a manual API key step entirely.
The specification's warning turned out to be correct. It cautioned that the federal research institute's careers page does not list actual vacancies. It doesn't.
One of the research documents was wrong. An agent implementing the same defense contractor's adapter found that the documented response shape didn't match reality. Jobs were nested differently, and fields lived inside an array instead of at the top level. Parsing the page according to the document returned zero jobs, a result that looks identical to a company simply having no openings.
Model behavior was measured, not assumed. Benchmarking local models rejected one outright. It spent its entire output budget reasoning and never closed its JSON. Testing also found that Ollama silently defaults to a 4,096 token context regardless of a model's advertised 262,000, which meant long job descriptions would have been classified from a fragment if nobody had checked.
The Bugs
Every bug below passed a green test suite. All of them were caught only by running against real sites with real data.
A completed run classified 500 jobs, but 413 of them had no description text attached. The model was handed only a title and a location, and confidently answered Neither for all of them, with reasoning that read no responsibilities provided. Because Neither is correctly left out of reports, the failure produced a plausible email saying four relevant jobs, with no error anywhere. It only came to light because the specification required the model to write down its reasoning, and that field became an audit trail. After the repair, the same run found 26 relevant jobs, including an AI/ML Engineer, an AI Evangelist, and a Large Language Model Specialist that had all been silently dismissed.
Four concurrent workers each held their own database connection, but SQLite's WAL mode allows exactly one writer at a time. Contention pushed past the timeout, and a perfectly healthy company got recorded as failed. Worse, the internal error was attributed to the career site itself, permanently adding to its failure count.
A single posting from a large aerospace and defense company, one out of 3,788, serialized without a title. The parser treated that as a structural failure and discarded the whole source rather than just the one record. The fix was a threshold rule: skip and count individual bad records, and only fail the source when essentially everything fails.
The code correctly avoided marking jobs as reported when a send failed, and a passing test proved it. But nothing downstream ever read that flag. Report selection was purely based on a time window, so any job stranded by one Gmail outage could never appear in a report again. The test checked that the flag was set. It never checked that the flag did anything.
Unit tests passed because they checked the plumbing, the data moving from database to email, not the actual text a person would read. The company name never made it into the template.
Auto detection could return the name of an adapter that was never registered, which silently created a source that would fail forever. The fix added a regression guard that fails the build if anyone adds a detection branch without a matching implementation. That guard immediately caught a second instance of the same bug.
A test used a hardcoded date instead of computing one relative to the real clock, so it quietly started failing as soon as the date it referenced was no longer today.
Repaired sources kept showing their old error message next to a green health badge. It is the mirror image of the first failure mode, and just as damaging: a source that looks broken when it isn't teaches people to stop reading the alerts.
What This Proves
- A green test suite is not evidence of a working product. Every defect above shipped through a passing suite, and every one was found by running against real data instead.
- The most dangerous failure is a confident, plausible wrong answer, not a crash. A crash is obvious. Four relevant jobs when the real number is twenty six is not.
- Requiring the model to explain its own reasoning built an audit trail that caught the worst bug in the project. It was a product requirement first, and became a debugging tool by accident.
- A broken source disguised as no jobs, and a healthy source flagged as broken, are equally corrosive. Both wear away trust in the signal.
- Delegation only works with clear ownership boundaries. Parallel agents worked safely because no two of them could write to the same file, so when infrastructure failures killed three at once, nothing was left corrupted.
Where It Falls Short
- If the Mac is fully powered off at the scheduled hour, that morning's run is skipped. launchd catches a run missed because the machine was asleep, not one missed because it was off, and the app detects and reports the miss.
- One of the eleven starter companies, a small technology company, has no machine readable job listings at all. It is marked unsupported rather than silently reporting zero jobs.
- One federal research laboratory's job portal is still unverified.
- Two of the four additional companies added through the interface needed manual configuration, because no recruiting platform sat behind the page. Pasting a URL works when a real ATS exists behind it. It cannot infer structure from a hand maintained web page, and a guessing heuristic was deliberately left out because it would fail more often than an honest error.
- Gmail send failure handling is covered by unit tests, not yet by an actual outage in production.
One More Detail
One company's careers page blocked the application outright while curl succeeded from the same machine in the same second. Cloudflare was fingerprinting the TLS handshake itself, not checking headers. An ordinary headless browser passed the challenge on its first attempt, with no evasion tooling involved. The fix shared the existing page parsing logic with a second retrieval strategy, so the next Cloudflare fronted career site is a configuration change rather than new code.
Built With
Curious what an agent team could build inside your business?
Hive Nova builds practical automation around the way your business actually works.
Book a Call