Summary
Every experiment here has a baseline, one change, a measurement plan written down before the results come in, and exactly one status. I measure by hand first, with a frozen set of prompts across six AI answer engines, and I publish the raw counts next to every verdict.
- Observed, Hypothesis and Conclusion are always labelled and kept apart.
- A page counts as Cited only when it’s a source in at least 2 of 3 runs.
- AI helps with research and drafts. It never makes the final call.
The big hypothesis
The whole site tests one claim: a developer with no marketing training can learn marketing, SEO and GEO with AI as tutor and practice partner, and get this blog indexed, ranked, cited by AI answer engines, and known. It is judged by dated markers, counted from launch day, and never by impression:
- 90 days: 10 to 15 pieces published, with search impressions showing in Google Search Console.
- 6 months, supported if: at least 2 of the 6 tracked engines cite this site for at least 1 non-branded prompt in the prompt set, and the site has at least 1 organic backlink.
- 12 months, refuted if: no tracked engine cites this site for any non-branded prompt.
How an experiment works
Each experiment gets an ID (EXP-001, EXP-002 and so on) and follows the same cycle:
- Observation: something I noticed, with the data that shows it.
- Question: what I want to know, narrow enough to test.
- Hypothesis: a prediction that could turn out wrong, written down before the change.
- Baseline: the numbers before anything changes.
- Intervention: the one change I make, and when.
- Measurement plan: what I measure, how, on which engines, and on which dates. It is fixed before the results come in.
- Results and interpretation: the numbers, then what I think they mean, including the limitations.
- Next hypothesis: what the result suggests testing next.
Every experiment carries exactly one status, always shown as a word and a symbol:
- – Planned: designed, with no measurement yet.
- ◐ Running: the baseline is taken and the measurement window is open.
- ✓ Supported: the results match the hypothesis.
- ✕ Refuted: the results contradict it.
- – Inconclusive: the data can’t tell either way. I say why.
What counts as evidence
- Observed is what I saw, with its source. Hypothesis is what I think explains it. Conclusion is what the evidence supports. I label each one, and a hypothesis never reads as a finding.
- Numbers come with their context: “12/40 prompts → 19/40 prompts (+58%)”, never a bare “+58%”. I always state the window and the engines covered.
- Correlation isn’t cause. With one site and no control group, most results here are observations, and I say so under Limitations.
- Advice from tools, blogs and AI is a hypothesis until I’ve tested it here.
How I measure
I measure by hand first, so I understand the mechanics before any tool does it for me.
AI answer engines
- Tracked engines: ChatGPT (with search), Perplexity, Google AI Mode, Gemini, Copilot and Claude (with web search). This list is fixed for the life of the big hypothesis, so results stay comparable. Google’s AI Overviews can’t be triggered on demand, so I read them from Search Console instead.
- Prompt set: a fixed list of prompts, frozen before the first measurement. Two are branded (they name GEO Geek or geo-geek.com) and six are topical, covering what this site wants to be found for. Each experiment publishes its prompts word for word.
- Runs: each prompt runs 3 times on each engine, in a fresh session. I’m logged out where the engine allows it, and use an account with no memory where it doesn’t.
- Cited means a geo-geek.com URL is among the answer’s sources in at least 2 of the 3 runs. I publish the raw count (for example 1/3) next to every verdict. Mentioned, meaning named without a link, is logged separately and never counts as Cited.
- An AI agent runs the prompts in a clean browser and saves every answer with a screenshot. I spot-check the saved answers.
Search
- Google Search Console: indexed pages, web impressions and clicks, and the AI Overviews report.
- Bing Webmaster Tools: indexed pages and AI citations.
- Server logs: the first date each search and AI crawler visits.
- Analytics: visits referred by AI answer engines.
site:search counts, which I treat as rough.
I don’t track rankings yet.
When a tool takes over
The first experiment’s verdicts use only the manual data. After its 90-day window, CitePulse, the GEO measurement product I help build, becomes the tracker for the later markers, running the same prompts on the same engines. Until then it runs privately alongside the manual measurement, and it doesn’t count toward any verdict. Where the two overlap, I keep both sets of numbers so they can be compared later.
How AI is used here
AI tutors, drafts and critiques. I decide, check and edit. Nothing goes live that I haven’t verified, and every post says what AI did in it. I mostly use Claude Code, for building this site and for writing it. Plenty of other AI tools will show up as subjects of experiments.
AI helps with: topic research, finding the questions people ask, first drafts of explanations, experiment ideas, outlines, title and summary variants, structured data and internal-linking suggestions, and summarizing results.
AI is not trusted alone with: final editorial judgment, facts that I haven’t checked, deciding what caused a result, or any decision that affects whether you can trust this site.
Corrections and changes
When later evidence changes a number or a conclusion, I correct the post in place, mark the correction at the point it applies, and update the post’s date. While an experiment is running, I log every change that reaches the live site, whatever prompted it, so the results can be read against it.