I built a crowd out of AI agents.
What can I test?
I had eight logos and nobody to ask.
So I built a crowd.
Each voter was a brand-new AI agent. It saw my brief and the options. Nothing else.
It picked one, said why, and was gone.
Fifty times over. Then I counted.

How it works
One brief. Up to three personas. Any options.
I wrote the brief in plain words.
I described each persona by their job. Never their taste.
Every agent became someone new. Random age, place, mood, rush.

Options could be anything. Text, images, PDFs, web pages. Or one sheet with every logo on it.

Agents never saw my file names or my order. final_best.png became opt-D-1.png. The order shuffled every call.

While it ran
Every vote was its own process.
New process. New session. No memory.

I could open any agent and see what it did. Only its private reasoning stayed hidden.

200 = 200
No web. No plugins. It could only open the test's own files.
Experiment 1
The test that fooled me.
Novexa, a medical billing company. Eight logos. Three people at a medical practice.
Fifty runs. $3.15.
That's list price. It runs on your own Claude account, so it's just token use. Free for you.
Every persona picked the same logo. All fifty times.
I thought the agents shared memory. They didn't.
My personas gave the answer away. “Wants friendly” picked the friendly one.
So I rewrote them. Job only. No taste.
Experiment 2
Novexa, done properly.

The brief: a cold letter from Novexa lands on a practice's desk. Which logo gets it opened?
A different person every run.
134 / 150

…looks like the letters we get from insurance companies or a law office.A 61-year-old receptionist, cardiology practice
Established beat clever.
Of the 16 who voted otherwise, 14 said it looked like mail from an insurer.
Experiment 3
I tested my own title screen.
My site's opening line, in eleven typefaces. Judged by the general public.


111 / 150

The top line looks kind of like a nice book or a café sign.A 36-year-old shop owner, rural West Virginia
What it isn't
A focus group.
These are simulated people. One model agrees with itself a lot.
A clear winner is a signal. So is a split.
A fast first read. Not the last word.
What I'll look at next:
Can AI testing stand in for a focus group, or will companies run both, side by side?
How many runs is enough?
Everyone will find their own use for AI.
In an age of abundance, this one helped me choose.
Run it yourself
Your turn.
The code is public. MIT.
Your machine. Your Claude account.
You need Node 22, Claude Code, and Chrome for web pages.
git clone https://github.com/ShaheryarWarraich/noorandmachine-public.git cd noorandmachine-public/test-with-ai/v1 npm start
Open 127.0.0.1:4330. Start with five runs.