898899 Lab
Testing Methodology

How We Test AI Tools

No AI-generated fluff. No marketing claims. Every tool runs through the same gauntlet.

The 5 Standardized Tasks

Every AI writing tool is tested with the same 5 prompts, designed to reflect real-world use cases:

Task 1: Blog Post Draft

Prompt: "Write a 500-word blog post explaining quantum computing to a marketing professional — no jargon, practical analogies."

Scored on: Accuracy, readability, structure.

Task 2: Email Sequence

Prompt: "Create a 3-email nurture sequence for a SaaS product launch. Include subject lines."

Scored on: Persuasiveness, personalization, format correctness.

Task 3: Social Media Copy

Prompt: "Write 5 LinkedIn posts promoting a new AI productivity tool. Vary the tone per post."

Scored on: Variety, hook quality, platform appropriateness.

Task 4: Product Description

Prompt: "Write a 150-word product description for noise-canceling headphones targeting audiophiles."

Scored on: Specificity, sensory detail, conversion potential.

Task 5: Data-to-Narrative

Prompt: "Turn these Q2 metrics [table provided] into a 3-paragraph executive summary with key takeaways."

Scored on: Data accuracy, insight quality, conciseness.

The 5 Scoring Dimensions

Each tool receives a 1-10 score on every dimension. Final score is a weighted average:

DimensionWeightWhat We Measure
Output Quality30%Accuracy, coherence, grammar, style adherence
Ease of Use20%UI intuitiveness, onboarding friction, feature discoverability
Speed15%Time from prompt to usable output (measured, not felt)
Features20%Templates, integrations, tone controls, team features
Value15%Price-to-quality ratio vs competitors at the same tier

Our Process

  1. 1 Manual Testing: A real human runs every task, records the session via OBS, and captures all output screenshots.
  2. 2 Structured Scoring: Results are logged into our JSON scoring template — no memory-based ratings.
  3. 3 AI-Assisted Draft: Raw data is fed into an LLM for a structured first draft, following our style guide.
  4. 4 Human Polish: Every draft gets manual review — fact-checking, de-AI-ification, screenshot insertion.
  5. 5 Continuous Updates: Tools are re-tested after major feature releases. Fresh data, always.