Transparency

How our exams are made

Every practice exam on Kwizza is AI-generated — we say that plainly, and we engineer around its weaknesses. Here is the exact pipeline, including the parts that exist purely to keep us honest.

  1. 1

    A blueprint before any questions

    We first derive the exam’s domain blueprint — the 6–16 topic areas the real exam actually covers, with their approximate weightings (for AWS Solutions Architect that means areas like compute, storage, networking, security). Every practice set for a subject is generated against this same blueprint, so coverage stays consistent and balanced instead of drifting set to set.

  2. 2

    Original questions, never copies

    Generation runs under strict originality instructions: never reproduce or paraphrase real exam questions, question banks, or “braindumps.” Each question must be a freshly invented scenario with its own names, numbers, and context. A per-set uniqueness seed pushes every generation run toward different phrasing, so two sets never converge on the same wording.

    This matters for two reasons: copying real exam content is a violation of certification NDAs and does not actually teach — and we want neither.

  3. 3

    Automatic duplicate sweeps

    After generation, every question is compared against every other question in the set and against all earlier sets for the same subject, using a text-similarity check. Anything too close to an existing question is dropped before you ever see it.

  4. 4

    Every source link is opened and checked

    Every question ships with a full explanation — why the right answer is right and why each distractor is wrong — plus links to official documentation (vendor docs, exam guides, standards bodies) where the concept is covered.

    Then we actually fetch every link before the question reaches you, and drop any that don’t resolve. This matters more than it sounds: when we first measured it, a large share of generated links were plausible-looking but dead — correct-looking paths on real sites that simply never existed. A model is good at the shape of a URL and bad at the path. Checking is the only thing that catches that, so we check.

  5. 5

    Automated item-writing checks

    Each question is then checked against published multiple-choice item-writing guidelines — the taxonomy from Haladyna, Downing & Rodriguez (2002) and the checkable rubric derived from it by Tarrant and colleagues (2006). These catch the flaws that let someone score well without knowing the material: the correct answer written longer and more carefully than the alternatives, a distinctive word shared only by the question and the key, “all of the above,” duplicated options.

    Questions with defects that make them unanswerable are removed. Style warnings are recorded and shown to you rather than hidden — you can see the counts for any set on its exam page.

  6. 6

    A provenance log for every set

    For each generated set we permanently record which model produced it, which version of our generation instructions was used, when it was created, and the provider’s request ID. That log exists so we can always demonstrate a set was independently created — and trace exactly how any question came to be.

  7. 7

    Real pass marks, honest numbers

    Where an exam has an official cut score (AWS ≈ 72%, CompTIA Security+ ≈ 75%), that’s the pass mark we show. We don’t inflate difficulty to flatter you or deflate it to sell you.

What AI-generated honestly means

AI models can make mistakes — occasionally a question is ambiguous or an explanation is imperfect. We’d rather tell you that than pretend otherwise. It’s why every question carries citations you can check, and why every question has a report button: flag it and a human reviews it.

Practice exams here are study material, not a replica of the real test. Used honestly — practice, review the explanations, chase the weak areas — they are one of the fastest ways to get ready.

What our numbers don’t tell you

Everything we report is something we measured while generating — links fetched, flags counted, topics covered. None of it is a score, and none of it has been certified by anyone. There is no accreditation available for practice exam material, and any study tool claiming one is worth a second look.

The two numbers that genuinely describe a question’s quality — how hard it actually is, and how well it separates people who know the material from people who don’t — cannot be calculated from the question text at all. They can only be computed once enough people have answered it. We don’t have that data yet, so we don’t claim those numbers. When we do, we’ll publish them here and retire the questions that fail.

See it for yourself — every exam in the library was made this way.

Browse the library