Visual AISKILLS
E-commercePromptsSkillsExplorer
Benchmark before comparison

AI product photography benchmark template for ecommerce.

Run a small product truth test before claiming an AI product photography tool is good enough for store, marketplace, landing page, or ad use.

Direct answer: A useful AI product photography benchmark tests multiple product categories, multiple scene types, and product truth failures, then records publish, repair, or reject decisions.
Test matrix

Test five product categories across three scene types.

Bottle or jar

Scenes: studio background, lifestyle surface, paid social crop

Failure risk: label drift, reflection drift, cap or lid deformation

Box or pouch

Scenes: marketplace image, seasonal background, landing page hero

Failure risk: package geometry, typography, fake badges

Transparent or glossy product

Scenes: white background, dark surface, shelf-style scene

Failure risk: transparency, material finish, contact shadow

Apparel or soft goods

Scenes: flat lay, model-free lifestyle, ad variant

Failure risk: fabric texture, scale, color and pattern drift

Accessory or small object

Scenes: macro crop, bundle image, comparison layout

Failure risk: scale, edges, missing parts, invented accessories

Scoring

Score product truth before judging visual quality.

Pretty outputs are not enough. Each image receives two review passes across identity, label, material, geometry, and channel claims before the team decides whether the output can be used commercially.

Identity and variant preserved
Label, logo, and claim text preserved
Material and color truth preserved
Geometry, scale, shadow, and contact points believable
No unsupported badges, discounts, ratings, endorsements, or platform logos
Evidence gallery

Sample evidence records

Use concrete sample records, not empty benchmark claims.

The template includes anonymized ecommerce sample records that show how product truth failures turn into publish, repair, or reject decisions. Treat them as a review-protocol demonstration, not a model ranking.

This is an anonymized review protocol demonstration, not a model ranking.

Sample record: VSK-BM-001

Amber supplement bottle

Prompt input: Premium PDP hero on a light stone surface

Product truth risk: Label drift and cap reflection distortion

Review gate: Check dosage text, cap width, label alignment, and bottle shape

Decision: Réparer
Score: 15/20

Sample record: VSK-BM-002

Matte coffee pouch

Prompt input: Warm kitchen background with seasonal props

Product truth risk: Invented badge and package geometry drift

Review gate: Check roast badge text, pouch corners, seal shape, and product count

Decision: Reject
Score: 10/20

Sample record: VSK-BM-003

Clear skincare serum

Prompt input: Dark reflective surface ad crop

Product truth risk: Transparency, liquid level, and contact shadow drift

Review gate: Check glass transparency, dropper metal, liquid level, and edge halos

Decision: Publish after final channel review
Score: 18/20

Sample record: VSK-BM-004

Sage cotton T-shirt

Prompt input: Model-free lifestyle flat lay with soft daylight

Product truth risk: Fabric color, collar rib, and scale drift

Review gate: Check source color, collar texture, fabric weight, and garment shape

Decision: Réparer
Score: 16/20

Sample record: VSK-BM-005

Measuring spoon set

Prompt input: Marketplace bundle image with clean background

Product truth risk: Included item count and size order drift

Review gate: Count each spoon, confirm nesting order, and reject invented accessories

Decision: Reject
Score: 12/20

Benchmark decision rules

18-20: publish after final channel review
14-17: repair failed details before publishing
0-13: reject or regenerate with narrower edit scope