How Your Submission Is Evaluated
AI Builder Series — 7 Quests #BuildInPublic Challenge
This document explains exactly how submissions to the challenge are scored. Nothing here is hidden or made up after the fact — every criterion below reflects either something taught across the program's 3 live sessions, or a requirement stated directly on the challenge page itself.
Scoring scale
Each quest you submit is scored out of 100 points:
- 40 points — Common criteria, identical across all 7 quests (below).
- 60 points — Task-specific criteria, different for each quest, listed further down.
If you completed more than one quest, each is scored independently on its own 100-point scale.
The 4 common criteria (apply to every quest)
| Criterion | Points | What we're looking for |
|---|---|---|
| Real, shareable, non-chat-bound artifact | 15 | Your submission is something we can actually open and use ourselves — a live URL, a downloadable file, a working repo/skill — not just a description of a chat conversation. |
| Iteration and judgment over one-shot prompting | 10 | Some evidence you tested the build, hit a rough edge, and refined it — rather than shipping the very first thing the AI produced, untouched. |
| Purposeful, discovery-minded build | 10 | A real reason is visible for why you built it the way you did — your own use case, your own product, your own angle — rather than a generic template with your name swapped in. |
| Responsible building practice | 5 | No leaked API keys/secrets visible anywhere in your repo, post, or screenshots; no spammy or scammy outreach where a quest involves contacting people or companies. |
Tracked, but does not affect your score
Whether you posted to LinkedIn with #BuildInPublic, tagged HelloPM, and showed the actual build (not just an announcement) is recorded — but it earns zero points. Posting is part of the challenge's participation rules, not part of how your work is graded. Your score is entirely about the quality of what you built.
Task-specific criteria, by quest
Quest 1 — Product Discovery interactive guide
"Build an interactive guide that teaches YOU Product Discovery."
| Criterion | Points | What good looks like |
|---|---|---|
| Actually interactive | 18 | Requires the learner to do something — answer a scenario, make a choice, get branched feedback — not just scroll static text or slides. |
| Content accuracy and depth | 24 | Reflects the discovery → delivery → distribution framing and the idea that small prototypes shown to users count as discovery, not just interviews/surveys. |
| Personalization | 18 | Some sign it teaches you specifically — your own product as the example, your own reflection notes — not a generic repurposed explainer. |
Quest 2 — Personal voice skill
"Create your personal voice skill — one that communicates in your style. Pull from LinkedIn posts and writing samples."
| Criterion | Points | What good looks like |
|---|---|---|
| Genuine personal source material | 18 | Actually seeded on your own LinkedIn posts/writing samples, not generic examples or someone else's writing. |
| Built as an actual reusable Skill, not a bare prompt | 24 | Packaged (name + description + instructions) so it can be reused across sessions/tools — not a one-off prompt typed into a chat once. |
| Demonstrated voice accuracy | 18 | Shows the skill's output next to a real post of yours so the match can be judged, not just taken on faith. |
Quest 3 — JD-to-resume customiser
"Build a tool that takes a job description and customises your resume to match. Outputs a finished PDF — not a chat-bound Gem."
| Criterion | Points | What good looks like |
|---|---|---|
| Outputs an actual finished PDF (hard requirement) | 24 | A real, downloadable, formatted PDF is produced. A Gem/GPT that only ever outputs text in a chat window scores 0 here regardless of text quality — the quest explicitly rules this out. |
| Genuine JD-to-resume matching | 24 | Responds to the specific JD's language and priorities — reordering/reweighting real experience — not cosmetic keyword-swapping. |
| Reusable as a tool | 12 | Works for more than one JD/resume pair as a repeatable pipeline, not a single manual edit done once. |
Quest 4 — Weekly funded-companies agent
"Build a weekly agent that scrapes last week's funded startups from the web and proposes your way in — a tailored application path."
| Criterion | Points | What good looks like |
|---|---|---|
| Real, recent funded-company data | 18 | Pulls genuine recent data from a real source — not a static or made-up list. |
| Genuinely tailored application path | 18 | Differs meaningfully per company — referencing the specific round, sector, or role — not the same paragraph with the name swapped in. |
| Recurring/automated structure | 18 | Some evidence of a scheduled or repeatable trigger — a workflow, cron job, or managed agent — matching the quest's "weekly" framing, not a single manual run. |
| No spammy outreach | 6 | The proposed outreach is something you'd genuinely do with judgment, not a mass cold-messaging script. |
Quest 5 — Duolingo-style AI concepts learner
"Build an app that makes you learn AI concepts the Duolingo way. Wire it up with PostHog for analytics."
| Criterion | Points | What good looks like |
|---|---|---|
| Genuinely gamified UX | 18 | Bite-sized lessons, a progress/streak/points mechanic, immediate feedback — not a static quiz or long-form article. |
| Teaches real AI concepts | 18 | Covers genuine AI/LLM vocabulary from the sessions (tokens, embeddings, attention, hallucination, etc.), not generic "what is AI" trivia. |
| PostHog actually integrated (hard requirement) | 24 | Real PostHog events fire from the app and a dashboard/event log can be shown as evidence. An app with no analytics wired up, or an unverifiable claim of integration, scores at or near 0 here. |
Quest 6 — GTM with AI-generated videos
"Run a full GTM with AI-generated videos (Higgsfield, Google Veo). Land your first 10 users on Instagram or LinkedIn."
| Criterion | Points | What good looks like |
|---|---|---|
| Actual AI-generated video content | 18 | Real video generated with an AI video tool — not a screen recording or stock footage. |
| Real GTM execution, not just a plan | 18 | The video was actually posted/run as a campaign — a plan describing what you would do does not satisfy this. |
| Evidence of real users acquired | 24 | Concrete evidence — signups, waitlist joins, trial users — tied to the campaign, ideally approaching the quest's 10-user bar. Likes/views alone are weaker evidence than real users. |
Quest 7 — Your own PM AI Agent
"Fork Hermes or OpenClaw on GitHub and ship your own Product Manager AI Agent."
| Criterion | Points | What good looks like |
|---|---|---|
| Actual fork of Hermes or OpenClaw (hard requirement) | 18 | A real, verifiable, linkable GitHub fork exists — not a from-scratch agent just described as "based on" one of them. |
| Meaningful PM-specific customization | 24 | Genuinely adapted toward a real PM use case — not an unmodified template or a single renamed system prompt. |
| Working, demonstrated agent | 18 | Shows the agent actually running and responding to a real PM task — not just a repo claimed to work but never shown operating. |
How we handle broken or incomplete submissions
We apply the same rule to everyone, so an unlucky link doesn't quietly cost you points with no explanation:
- A link that doesn't open, 404s, or is private is scored 0 with an explicit note — never
silently skipped.
- A LinkedIn post announcing intent with no actual artifact behind it scores 0 on the
real-artifact criterion specifically, with a note.
- A link that clearly matches a different quest than claimed is re-scored against the quest it
actually matches, if identifiable, with a note explaining the reassignment.
- Partially accessible submissions (e.g. a screenshot works but the live link doesn't) are
graded only on what we can independently verify — unverifiable claims in a caption don't earn credit, but everything we could confirm still counts.
- Every 0, partial, or capped score always comes with a one-line reason — an unexplained zero is
treated as a mistake in our process, not a valid outcome.
How the grading actually happens
For each submission, we open the link ourselves — for a live app, we click through the actual flow rather than assuming it works; for a repo or file, we open what's inside. We score each criterion in the tables above individually and write a short justification tied to what we actually observed, not a generic impression. Where something is unverifiable (a claim with no evidence to check), we score conservatively and note exactly what couldn't be confirmed, rather than giving the benefit of the doubt either way.