Skip to main content

Blog

Founder Hub

Your AI Coding Benchmark Score Is Not Your Software Quality

· 16 min read
Codalio Team
AI app builder team

Over the past year, I have spent more time than I expected looking at AI coding benchmarks. Every few weeks, it seems, a new model arrives with another impressive result. SWE-bench scores go up, leaderboards change, and for a brief period one model appears to have established a meaningful lead—until the next release changes the picture again. As someone involved in building software products, I understand why we pay attention. If I am going to put an AI model somewhere inside a software development process, I want evidence that it can actually perform software engineering tasks.

What Canadian Founders Need Before Applying to ventureLAB's Hardware Catalyst

· 6 min read
Codalio Team
AI app builder team

What does ventureLAB's Hardware Catalyst actually ask founders for?​

Before you apply to ventureLAB's Hardware Catalyst Initiative, you need three things in writing: a product plan, an IP strategy, and a customer-engagement model. Together, those three documents make up a product spec. Founders who write the spec first can build the application in a day. Founders who skip it can spend a month on it.

Codalio Blueprint — Stop Agents From Building the Wrong Thing

· 5 min read
Codalio Team
AI app builder team

Codalio Blueprint is a free, MIT-licensed plugin that adds 10 planning and review skills to popular coding agents. The flagship skill, prd-builder, runs three lenses (Product & Scope, Architecture & Data, GTM) and synthesizes a single PRD that resolves contradictions before anyone writes code. Later skills review what gets built — auth exposure, performance cost, silent regressions, and which tests are worth writing. Outputs land as files in your repo — not ephemeral chat — so builders read a durable source of truth.

Codalio Blueprint hero — 10 planning skills, one install, zero code until you approve

Best Open-Source Coding Skills, Plugins & AI Agents (Updated Weekly)

· 10 min read
Codalio Team
AI app builder team

Last updated: September 11, 2026. Reviewed weekly.

If you're picking three open-source coding tools today: BMAD-METHOD to plan before you build (the only planning tool here that runs in a browser with no terminal), Goose to do the building (the only mainstream coding agent with a real desktop app), and the Claude Code GitHub Action to review what the agent wrote. All three are OSI-licensed and all three were pushed to within the last 48 hours.

This page is written for founders who can't code and are building anyway. Every star count below was read on September 11, 2026 — the first complete audit this page has had. The last two runs were cut short by GitHub's rate limit; there's now a way around it, explained at the end.

Last week's four ownership changes have all held: Goose is settled at aaif-goose, OpenCode at anomalyco, OpenHands at OpenHands, PR-Agent at The-PR-Agent. All four were pushed to again this week, so the transfers look like housekeeping rather than abandonment.


Idea to plan​

Before you write code you need something to build against. This is the category most vibe-coded projects skip, and skipping it is why they stall at 70%.

GitHub Spec Kit — 135,587 stars, MIT. Turns an idea into constitution, spec, plan and tasks across 30+ agents. The most rigorous option, and it opens with uv tool install and a Python 3.11 requirement.

OpenSpec — 67,989 stars, MIT. Proposals, specs and task checklists before coding, with a local dashboard.

BMAD-METHOD — 52,900 stars, licence unresolved (see below). Runs agile agent roles from idea to working software, with ChatGPT and Gemini web bundles.

Task Master — 28,063 stars, licence unresolved, last pushed April 28. Breaks a PRD into ordered, dependency-aware tasks.

Backlog.md — 6,704 stars, MIT. A markdown task board inside your git repo, with a local kanban UI.

codalio-blueprint — 5 stars, MIT. This one is ours. Turns a rough idea into a written PRD.

The one we'd install first: BMAD-METHOD. It's the only planning tool here with a genuine no-terminal on-ramp — the web bundles run as ChatGPT Custom GPTs and Gemini Gems, so you can do the entire planning phase before installing anything. Spec Kit is more rigorous and has two and a half times the stars, but its install command loses exactly the reader this page is for. The honest limitation on BMAD: the web bundles cover planning only. The moment you start implementing you're back in a CLI.

BMAD's licence is still unresolved, two weeks on. The README badge says MIT; GitHub's detector still returns no assertion, which normally means the LICENSE file has been edited. It's still our top pick and we're still not claiming the licence changed — but two weeks is long enough that this isn't a blip. If you're bundling BMAD into something you sell, open the LICENSE file and read it yourself. We've changed the licence column from "MIT" to "unresolved" to stop implying a grant we can't verify. Same for Task Master.

Task Master is the one we'd now hesitate over. Unresolved licence, and last pushed April 28 — four and a half months, unchanged from last week. Inside our six-month window, so it stays, but it comes off the page at the end of October if nothing lands.

On our own tool, plainly. codalio-blueprint runs three lenses (Product & Scope, Architecture, GTM) and synthesizes one PRD rather than stapling three documents together. We think that's well-built. It's also five weeks old, has 5 stars and 2 forks — unchanged from last week — no external contributors, and no test of whether the PRD is any good beyond examples we wrote ourselves. Last week we reported three new stars. This week, none. That's what a flat week looks like. Only Claude Code has real install instructions, whatever the README implies. Don't pick it over BMAD on our say-so.

Building​

Superpowers (285,156, MIT) installs a full agent methodology as composable skills across ~14 hosts. mattpocock/skills (259,497, MIT) applies senior-engineer review and TDD. OpenCode (206,685, MIT) runs a terminal coding agent against any model provider. Anthropic Skills (175,796, no root licence) holds the official reference skills and spec. OpenAI Codex CLI (123,351, Apache-2.0) and Gemini CLI (106,917, Apache-2.0) both run local coding agents. OpenHands (87,418, MIT) gives an agent a browser, terminal and editor. Cline (67,832, Apache-2.0) plans then edits with approval steps. Context7 (61,882, MIT) feeds agents version-correct library docs. Goose (54,129, Apache-2.0) runs an autonomous agent with a desktop app. Continue (35,869), vercel-labs/skills (31,395), Serena (29,183), Vibe Kanban (28,055) and Kilo Code (27,262) round it out.

The one we'd install first: Goose. The only mainstream open-source coding agent with a real desktop application — you see a window instead of a terminal — and Apache-2.0 with any-LLM support means no lock-in to one vendor's pricing. OpenCode has nearly four times the stars and is better if you're comfortable in a terminal, but it assumes you already are. Honest limitation: the desktop app hides the terminal, not the concepts. Extensions and MCP configuration still expect developer vocabulary, and you'll hit that wall on day two.

Goose has been pushed to repeatedly since moving to aaif-goose, the licence is unchanged, and it's up 230 stars on the week. One week isn't a guarantee, but it's the evidence we said we'd go and look for.

Vibe Kanban hasn't been pushed to since April 24 — same clock as Task Master.

Two things worth knowing before you install from this group. "Open source" often means the wrapper, not the engine: Codex CLI is Apache-2.0 and useless without a paid OpenAI plan, and Gemini CLI's free tier is a Google account benefit that can change without the repo changing. And the Anthropic Skills repo still has no root LICENSE file — confirmed again this week. Licensing is per-skill, and the document skills are source-available rather than open source.

Reviewing and QA​

This is where non-technical founders are most exposed. An AI wrote your code; something other than the same AI should look at it.

Trivy (37,870, Apache-2.0) scans dependencies, containers and IaC. Playwright MCP (37,011, Apache-2.0) lets an agent click through your live app. Gitleaks (29,234, MIT) detects committed secrets. Semgrep (16,590, LGPL-2.1) scans for security bugs. PR-Agent (12,950, MIT) reviews pull requests. Claude Code GitHub Action (8,844, MIT) reviews when you mention @claude. Trail of Bits Skills (7,041, CC-BY-SA-4.0) adds professional audit skills. cc-safety-net (1,535, MIT) blocks destructive commands.

The one we'd install first: the Claude Code GitHub Action. Typing "@claude review this" on a pull request is the lowest-literacy way to get a real second opinion on agent-written code, and it's MIT with no paid tier. Pair it with Gitleaks — an AI reviewer will happily discuss your architecture while ignoring the API key you committed in week one. Honest limitation: free to install, not free to run. Every review burns API credits and there's no built-in spend cap.

PR-Agent looks healthy after last week's move out of the Qodo org — pushed to this week, up about a hundred stars. The worry we raised hasn't materialised.

Trail of Bits Skills are the real thing, written by an actual security firm, but CC-BY-SA-4.0 is a content licence with a share-alike obligation. Read it before bundling commercially.

Shipping​

Supabase (109,056, Apache-2.0), Docusaurus (66,226, MIT), Coolify (61,677, Apache-2.0), GitHub MCP Server (32,869, MIT), semantic-release (24,034, MIT) and Changesets (12,384, MIT). All six were re-read on September 11 — these are the entries that carried stale August numbers for two weeks — and all six were pushed to within the last four days.

The one we'd install first: Supabase. The one piece of shipping infrastructure a non-technical founder can genuinely operate alone: clickable console, real free tier, Apache-2.0 so you can leave with your data. Coolify is better once you outgrow it, but the hardest step happens before Coolify appears — you have to rent a VPS and SSH into it. Honest limitation: some hosted-platform pieces aren't in the Apache-2.0 repo, so self-hosting isn't feature-equivalent.

One warning on the GitHub MCP Server: it needs a personal access token, and the easy broad-scope token hands an agent write access to every repository you own. Scope it down.

What didn't make the list​

Aider — still maintained, still excellent, wrong for this audience: its whole interaction model assumes you think in git commits and diffs. 48,897 stars, Apache-2.0, last pushed May 22 (our first current read on it). That date is approaching four months, which is worth watching on a tool this widely recommended.

Dokploy — open-core presented as open source, and the API confirms it: no licence assertion. Apache-2.0 applies only outside a /proprietary directory, and the proprietary licence forbids production use without a commercial agreement. Coolify is genuinely Apache-2.0 throughout and gets the slot.

Qodo-Cover — abandoned, with an explicit "no longer maintained" notice; the successor is paid. Automated test generation remains a real hole with no good open-source answer.

gpt-engineer, Devika, Claudia, snarktank/ai-dev-tasks, coderabbitai/ai-pr-reviewer — dead, stale, or 404. Named rather than silently omitted, because several still rank near the top of listicles on star count alone.

How to get exact star counts without hitting GitHub's rate limit​

Worth sharing, because it broke this page's audit twice. GitHub's unauthenticated REST API allows 60 requests an hour, which doesn't cover a 35-repo page, let alone four pages. But every repository page embeds its own exact figure in the HTML as "stargazerCount": <n> — the same number the API returns, not the rounded "48.9k" the page displays. Reading that costs no API quota. That's how every figure here got a current date for the first time.

Frequently asked questions​

How often is this updated? Every week. Entries we cannot verify are removed or flagged rather than quietly kept.

Is this the full page? Yes — this page is the canonical living guide, and it is updated here every week.

What changed this week​

Fixed the thing that kept breaking: every star count is now current. All 35 entries read September 11, including the nine that had carried August 28 numbers for two consecutive runs.

Corrected our own entry, in the unflattering direction. codalio-blueprint is still at 5 stars and 2 forks — flat on the week, not rising as we implied last week. A spot-check earlier in this run misread it as 1 star; 5 is correct. We'd rather print the correction than let either number stand.

Hardened two licence columns. BMAD-METHOD and Task Master now read "unresolved" rather than "MIT" — neither resolves to a standard SPDX licence, two weeks running. We're not claiming either changed; we're refusing to keep printing a grant we can't verify.

Maintenance clocks, now dated: Task Master last pushed April 28, Vibe Kanban April 24. Both stay this month, both come off at the end of October if nothing lands.

Last week's ownership changes all look healthy — Goose, OpenCode, OpenHands and PR-Agent each pushed to this week under their new owners, licences unchanged.

Notable movers: mattpocock/skills +10,881, superpowers +3,554, OpenCode +2,995, spec-kit +2,177, codex +1,923, OpenHands +1,270.

Nothing incomplete this run. For the first time since this page launched, there's no "we couldn't verify this" list.

Best Open-Source Fundraising Skills, Plugins & AI Agents (Updated Weekly)

· 9 min read
Codalio Team
AI app builder team

Last updated: September 11, 2026. Reviewed weekly.

If you're picking three open-source fundraising tools today: Papermark as your data room (AGPL-3.0, the most battle-tested DocSend replacement there is), Outreachr as your investor CRM (Apache-2.0, local-first, actual software rather than a prompt pack), and nock to pressure-test your deck before an investor does it for you.

One honest framing before the list. Fundraising is the thinnest of the four open-source categories we track, and it's thin in a specific way: deck-prep skills are abundant and mostly interchangeable, while investor data — the thing that would actually save you a week — is almost entirely locked behind paid APIs. We mark where the shelf is empty rather than padding it.

Timing note, because it's the reason you're reading this. Y Combinator's Winter 2027 application window closes November 2, 2026 at 8pm PT — seven weeks. Techstars' synchronized Spring 2027 deadline is November 18. Both are on our events, grants and fundraising deadlines page, verified the same day as this one: https://codalio.com/blog/startup-events-grants-fundraising-deadlines

Every star count here is current as of September 11, 2026 — the first complete read this page has had. It settles a number we flagged as disputed last week, and the answer is not the one we guessed.


Deck prep​

Lenny Skills — 1,319 stars, MIT structure. 76 PM and founder skills including fundraising and exits playbooks.

ppt-agent-skills — 893 stars, no licence shown. Generates PPTX decks from prompts, with roadshow templates.

claude-skills-founder (89, MIT), fluiddocs-deck-builder (30, MIT), vc-skills (30, MIT — simulates 28 named VCs), cc-skills-vc-fundraising (24, MIT), nock (10, MIT), founder-skills (9, MIT), pitch-deck-mastery-skill (8, MIT), startup-problem-finder (6, MIT).

OpenStartupModel's cap table is a free spreadsheet with no licence at all.

The one we'd install first: nock. Every other tool in this group writes slides. nock tells you which slide will get you killed — it's built from one seed investor's questions across 53 real pitch and diligence meetings. Fixing a deck is easy; knowing what's actually broken in it is not. Honest limitation: it's one investor's lens, narrow and idiosyncratic by design. Run it and then run vc-skills for a second opinion, and remember that vc-skills' "firm personalities" are the author's characterisation, not sourced from those firms.

New flag: ppt-agent-skills shows no licence. At 893 stars it's the second-largest entry in this group, and GitHub resolves no licence for it. Last week we printed "MIT" on the strength of its README. It generates decks you might send to an investor, so read the repository yourself before relying on it commercially.

Read the small print on two more. OpenStartupModel calls itself open source and carries no licence — it's a law firm's goodwill marketing, marked "educational purposes only." And cc-skills-vc-fundraising and claude-skills-founder each have four commits and describe themselves as grounded in Sequoia and a16z frameworks; that's the author's synthesis, uncited. Useful scaffolding, not authority. Both grew this week, which changes nothing about that.

Investor research​

This bucket is genuinely thin, and that's the most useful thing on the page.

awesome-oss-investors (473, MIT) lists 80+ VCs investing in commercial open source with ticket sizes — genuinely good, but last committed well over a year ago, so ticket sizes and fund status will have drifted. openbook (64, MIT, last pushed August 1) scrapes and publishes an open VC database. awesome-startup-fundraising (7, MIT) is vendor-curated. OpenVC is proprietary with a free tier.

The one we'd start with: openbook. The only genuinely open, actively maintained attempt at a public VC database rather than a static list or a freemium funnel. The honest limitation is big: the repository holds scrapers, not rows. The dataset lives on DoltHub and we couldn't render that page to confirm its size or freshness, so treat this as a tool for building your own list, not a list.

For most founders the pragmatic answer is OpenVC's free tier — 16,000+ investor profiles, filterable by thesis, exportable. It isn't open source and we're saying so rather than smuggling it into an open-source list.

What's missing: there is no open-source thesis-matching agent. Nothing takes your company and returns a ranked, reasoned investor list. Every MCP server in this space — Crunchbase, Affinity, Harmonic — is a thin client for a paid proprietary API.

Outreach​

ECC — 256,281 stars, MIT. A 286-skill, 68-agent harness; the fundraising pieces are two skills inside it. A heavy install for a small payload unless you were adopting the whole thing anyway.

Outreachr — 259 stars, Apache-2.0, v0.1.1, 41 commits. Local-first investor CRM: relationship mapping, warm-intro planning, email approval workflow, pipeline.

venture-ops (7, MIT) and founder-fundraising-outreach (5, MIT).

The one we'd install first: Outreachr. The only tool in this category that's an actual investor CRM rather than instructions for writing emails. Apache-2.0 and local-first, with investor data in a local SQLite vault and encrypted secrets. Honest limitation: v0.1.1, about six weeks old, and the macOS and Windows builds are unsigned and unnotarised — you'll click through security warnings to install something that will hold your investor pipeline.

Last week's disputed number is settled, and we were wrong about which figure was wrong. We recorded 238 stars from the API on August 28, then a hand-check on September 4 appeared to show 6 stars and 10 commits, so we marked it disputed and told you you'd be adopting a six-star project. The correct figure is 259 stars and 41 commits. The API was right all along; the hand-check was the error — almost certainly a misread of a partially-rendered page, which is exactly the failure mode we warn about elsewhere when we tell you not to trust rendered GitHub counts.

We're leaving the history visible rather than quietly swapping the number, because the correction cuts against us twice: we published a wrong figure, and we published it in the direction that made a tool we recommend look weaker than it is. If you skipped Outreachr last week because this page called it a six-star project with ten commits, that was our mistake. Still young, still unsigned, still a real bus-factor risk — but not the thing we described.

The underlying argument hasn't changed: in a category this small, low adoption is information rather than disqualification. A genuinely useful tool here might have 10 stars where a coding tool has 100,000.

founder-fundraising-outreach is a competitor's skill and it's real: MIT, self-contained, eight distinct email modes including stall recovery, which is the mode most founders actually need and never plan for. It's also three commits old and tagged v0.9 pilot. We'd rather list it honestly than pretend it doesn't exist.

Data room and diligence prep​

Papermark — 9,158 stars, AGPL-3.0, 5,000+ commits. Self-hosted data room with per-page view analytics.

due-diligence-portal — 1 star, Apache-2.0. A single Docker container with NDA gating and an audit log. Genuinely nice shape for a small raise. You are the QA team.

YC SAFE documents — free, no stated licence, the closest thing to a default in pre-seed financing.

The one we'd install first: Papermark. By a wide margin the most battle-tested open-source DocSend replacement. The per-viewer page analytics are exactly what you want during a raise — knowing which investor spent four minutes on your financials changes your follow-up. Honest limitation: self-hosting means Postgres, blob storage and SMTP wiring. Budget an afternoon of infrastructure, not a one-click deploy. There's a paid hosted tier if you'd rather not.

Papermark's move has completed cleanly: mfts/papermark redirects to papermark/papermark, both report the same 9,158 stars, licence unchanged. Update any pinned remote.

On the YC SAFE: free, standard since the post-money version in 2018. It carries no licence grant and it is not a substitute for a lawyer — it's a starting document that saves your lawyer time. If you're applying to Winter 2027, read it before November 2, not after.

What didn't make the list​

captableinc/captable — self-hostable cap table management, last pushed June 2025 with cap-table management itself still marked work in progress. Over a year stale on something this consequential isn't "stable", it's abandoned in place. selcuke/venture-capital-firms-list — last pushed May 2019, no licence. Seven-year-old VC contact data is worse than none. Graphite Financial's "open source financial model" — no licence anywhere and the download gated behind a form. That's a lead magnet. Open-Term-Sheet / SAFE-Note-Translations — five-year-old legal documents are actively dangerous, not merely stale. Crunchbase MCP server — a thin client for a paid key, so no free path.

Frequently asked questions​

How often is this updated? Every week. Entries we cannot verify are removed or flagged rather than quietly kept.

Is this the full page? Yes — this page is the canonical living guide, and it is updated here every week.

What changed this week​

Settled the Outreachr dispute, against ourselves. 259 stars, 41 commits, v0.1.1. The August 28 API reading was substantially right; the September 4 hand-check was the error. We told readers they'd be adopting a six-star project — that was wrong, and if it put you off a tool we otherwise recommend, the mistake was ours.

New licence flag: ppt-agent-skills (893 stars) shows no licence. We printed "MIT" last week on the strength of its README. GitHub resolves nothing.

First complete star audit this page has had. All 19 GitHub entries read September 11. Notable: Lenny Skills 1,287 → 1,319, claude-skills-founder 66 → 89, ECC 243,927 → 256,281, Papermark 8,980 → 9,158.

Papermark's move confirmed complete. Old path redirects, identical figures, AGPL-3.0 unchanged.

Timing updated: YC Winter 2027 is now seven weeks out. Techstars Spring 2027 is November 18.

Where the shelf is still empty: no open-source thesis-matching agent; no free path to investor data; the best cap-table tool remains stale. Unchanged, and unlikely to change soon.