HelloPM · AI Builder Series Explore builds →

How Your Submission Is Evaluated

AI Builder Series — 7 Quests #BuildInPublic Challenge

This document explains exactly how submissions to the challenge are scored. Nothing here is hidden or made up after the fact — every criterion below reflects either something taught across the program's 3 live sessions, or a requirement stated directly on the challenge page itself.

Scoring scale

Each quest you submit is scored out of 100 points:

If you completed more than one quest, each is scored independently on its own 100-point scale.

The 4 common criteria (apply to every quest)

CriterionPointsWhat we're looking for
Real, shareable, non-chat-bound artifact15Your submission is something we can actually open and use ourselves — a live URL, a downloadable file, a working repo/skill — not just a description of a chat conversation.
Iteration and judgment over one-shot prompting10Some evidence you tested the build, hit a rough edge, and refined it — rather than shipping the very first thing the AI produced, untouched.
Purposeful, discovery-minded build10A real reason is visible for why you built it the way you did — your own use case, your own product, your own angle — rather than a generic template with your name swapped in.
Responsible building practice5No leaked API keys/secrets visible anywhere in your repo, post, or screenshots; no spammy or scammy outreach where a quest involves contacting people or companies.

Tracked, but does not affect your score

Whether you posted to LinkedIn with #BuildInPublic, tagged HelloPM, and showed the actual build (not just an announcement) is recorded — but it earns zero points. Posting is part of the challenge's participation rules, not part of how your work is graded. Your score is entirely about the quality of what you built.

Task-specific criteria, by quest

Quest 1 — Product Discovery interactive guide

"Build an interactive guide that teaches YOU Product Discovery."

CriterionPointsWhat good looks like
Actually interactive18Requires the learner to do something — answer a scenario, make a choice, get branched feedback — not just scroll static text or slides.
Content accuracy and depth24Reflects the discovery → delivery → distribution framing and the idea that small prototypes shown to users count as discovery, not just interviews/surveys.
Personalization18Some sign it teaches you specifically — your own product as the example, your own reflection notes — not a generic repurposed explainer.

Quest 2 — Personal voice skill

"Create your personal voice skill — one that communicates in your style. Pull from LinkedIn posts and writing samples."

CriterionPointsWhat good looks like
Genuine personal source material18Actually seeded on your own LinkedIn posts/writing samples, not generic examples or someone else's writing.
Built as an actual reusable Skill, not a bare prompt24Packaged (name + description + instructions) so it can be reused across sessions/tools — not a one-off prompt typed into a chat once.
Demonstrated voice accuracy18Shows the skill's output next to a real post of yours so the match can be judged, not just taken on faith.

Quest 3 — JD-to-resume customiser

"Build a tool that takes a job description and customises your resume to match. Outputs a finished PDF — not a chat-bound Gem."

CriterionPointsWhat good looks like
Outputs an actual finished PDF (hard requirement)24A real, downloadable, formatted PDF is produced. A Gem/GPT that only ever outputs text in a chat window scores 0 here regardless of text quality — the quest explicitly rules this out.
Genuine JD-to-resume matching24Responds to the specific JD's language and priorities — reordering/reweighting real experience — not cosmetic keyword-swapping.
Reusable as a tool12Works for more than one JD/resume pair as a repeatable pipeline, not a single manual edit done once.

Quest 4 — Weekly funded-companies agent

"Build a weekly agent that scrapes last week's funded startups from the web and proposes your way in — a tailored application path."

CriterionPointsWhat good looks like
Real, recent funded-company data18Pulls genuine recent data from a real source — not a static or made-up list.
Genuinely tailored application path18Differs meaningfully per company — referencing the specific round, sector, or role — not the same paragraph with the name swapped in.
Recurring/automated structure18Some evidence of a scheduled or repeatable trigger — a workflow, cron job, or managed agent — matching the quest's "weekly" framing, not a single manual run.
No spammy outreach6The proposed outreach is something you'd genuinely do with judgment, not a mass cold-messaging script.

Quest 5 — Duolingo-style AI concepts learner

"Build an app that makes you learn AI concepts the Duolingo way. Wire it up with PostHog for analytics."

CriterionPointsWhat good looks like
Genuinely gamified UX18Bite-sized lessons, a progress/streak/points mechanic, immediate feedback — not a static quiz or long-form article.
Teaches real AI concepts18Covers genuine AI/LLM vocabulary from the sessions (tokens, embeddings, attention, hallucination, etc.), not generic "what is AI" trivia.
PostHog actually integrated (hard requirement)24Real PostHog events fire from the app and a dashboard/event log can be shown as evidence. An app with no analytics wired up, or an unverifiable claim of integration, scores at or near 0 here.

Quest 6 — GTM with AI-generated videos

"Run a full GTM with AI-generated videos (Higgsfield, Google Veo). Land your first 10 users on Instagram or LinkedIn."

CriterionPointsWhat good looks like
Actual AI-generated video content18Real video generated with an AI video tool — not a screen recording or stock footage.
Real GTM execution, not just a plan18The video was actually posted/run as a campaign — a plan describing what you would do does not satisfy this.
Evidence of real users acquired24Concrete evidence — signups, waitlist joins, trial users — tied to the campaign, ideally approaching the quest's 10-user bar. Likes/views alone are weaker evidence than real users.

Quest 7 — Your own PM AI Agent

"Fork Hermes or OpenClaw on GitHub and ship your own Product Manager AI Agent."

CriterionPointsWhat good looks like
Actual fork of Hermes or OpenClaw (hard requirement)18A real, verifiable, linkable GitHub fork exists — not a from-scratch agent just described as "based on" one of them.
Meaningful PM-specific customization24Genuinely adapted toward a real PM use case — not an unmodified template or a single renamed system prompt.
Working, demonstrated agent18Shows the agent actually running and responding to a real PM task — not just a repo claimed to work but never shown operating.

How we handle broken or incomplete submissions

We apply the same rule to everyone, so an unlucky link doesn't quietly cost you points with no explanation:

silently skipped.

real-artifact criterion specifically, with a note.

actually matches, if identifiable, with a note explaining the reassignment.

graded only on what we can independently verify — unverifiable claims in a caption don't earn credit, but everything we could confirm still counts.

treated as a mistake in our process, not a valid outcome.

How the grading actually happens

For each submission, we open the link ourselves — for a live app, we click through the actual flow rather than assuming it works; for a repo or file, we open what's inside. We score each criterion in the tables above individually and write a short justification tied to what we actually observed, not a generic impression. Where something is unverifiable (a claim with no evidence to check), we score conservatively and note exactly what couldn't be confirmed, rather than giving the benefit of the doubt either way.