Logged out. Real. Not staged.
Verbatim. Flowcraft never comes up.
Flowcraft shows up on 2 of the 24 pages assistants actually cite for this category.
51% of B2B buyers now start research in an AI chatbot. 92% say it shapes their shortlist. (G2, Mar 2026)
This is the full gap list: 33 findings across nine surfaces, plus what to do about each one and how you'd check it worked.
Where Flowcraft stands in AI-generated shortlists for "project management software". Every gap we found, what each one costs, and what changes if this runs continuously instead of once.
Share of AI-cited category pages that name Flowcraft, weighted by how often each page is quoted. Category leader sits at 71.
Across ten surfaces, each with named evidence and a fix.
The uncomfortable part isn't the score. It's that 13 of the 24 pages an assistant cites for this category are either yours to edit or open to submission today, and none of the 13 currently carry Flowcraft. Those are forms nobody has filled in.
14 vendors named: Taskforge, Basecraft, Workstream, TaskHive, ProjectPilot, Boardwise and others. Flowcraft wasn't one of them.
"One of the leading horizontal enterprise AI agent platforms." 4-5★ on capability and integrations. Closest comparison: TaskHive.
Sixteen ways of asking the same commercial question. A filled cell means the vendor was named in that answer.
Worth noting what this rules out: the shortlist isn't randomly reshuffling. The same five to seven names hold across almost every phrasing, which means the set is stable enough to be worth entering, and stable enough that entering it holds.
Both of Flowcraft's mentions are weak: unrated on Gartner Peer Insights, and #7 on one listicle. Flowcraft Sprints, the flagship product, appears on none of the six time-tracking-specific lists. The sub-category where Flowcraft has the strongest product claim is the one where it is most completely absent.
Brand mentions like these correlate with AI-answer visibility three times more strongly than backlinks do: r=0.664 vs. 0.218 across a 75,000-brand study (Ahrefs). That's the reasoning behind chasing mentions here rather than link-building.
Most teams treat AI visibility as a black box. It isn't. Every answer is assembled from a countable set of pages, and each of those pages has an owner.
Listicles and directories that accept vendor submissions today, free or paid. Flowcraft is on 1 of the 9. No gatekeeper, no relationship required. A form and a follow-up.
G2, Capterra, Gartner Peer Insights, Crunchbase. Flowcraft controls the text on all four. Three currently contain errors or omissions that a model can quote back.
Independent analyst posts and trade coverage. Not purchasable; it needs an angle and outreach. This is the slow half of the plan.
Pages published by TaskHive, Workstream and ProjectPilot that rank for category and "alternatives" queries. Flowcraft will never appear on these, but they show exactly which play is working.
The same 24 pages, sorted a second way: by what they argue. Ownership says who can change a page. Stance says what a model repeats after reading it. A page that argues against you is not the same problem as a page that omits you, and it does not get fixed the same way.
| Stance toward Flowcraft | Pages | Why it is scored this way |
|---|---|---|
| None hostile | 0 | No cited page currently argues against Flowcraft by name. This is the one genuinely good number in the report, and it is the number most likely to change without warning. |
| Named, weak | 2 | The two pages that contain the string "Flowcraft" at all: an unrated Gartner Peer Insights listing and a #7 listicle placement. Neither is negative. Neither is worth much either. |
| Against, unnamed | 4 | Competitor-owned pages arguing the category in terms that exclude Flowcraft, without ever printing the name. A monitor that only searches for "Flowcraft" scores these as absence. They are not absence. |
| No engagement | 18 | Neither for nor against. These are the pages the submission and profile work in Act II is aimed at. |
These are the pages that answer "who is this vendor" when a model needs to check. Each is either working for Flowcraft, working against it, or missing.
What gets cited for "best X" phrasings, by source type. The black slice is the one Flowcraft has a single placement on.
Most of these take submissions. This is the cheapest gap in the report to close. Two details that change the priority:
Position matters more than presence. Models quote the top of a list far more often than the tail. The single placement Flowcraft does hold sits at #7, below where extraction usually stops.
Freshness decays. 79% of the lists ChatGPT cites for this kind of question were updated within the last year. Flowcraft's one placement is on a page last refreshed 14 months ago, which puts it on the wrong side of that line and falling.
Before any of the above matters, the machines have to be able to fetch and parse the site. Most teams have never audited this list. Several of these agents did not exist eighteen months ago.
One thing deliberately not on the fix list: llms.txt. 97% of published llms.txt files are never fetched by any AI crawler, and none of the three engines here use it as a ranking signal. It is the most-recommended and least-useful item in this space, and adding it would be theatre.
Underneath retrieval sits an entity layer: the model's internal answer to "who is this, and what do I already know about them." A fragmented entity means every signal counts for less than it should.
Same eighteen profiles on both sides. The only difference is whether the graph knows they describe one company.
sameAs markup tying the profiles together and no Wikidata record to anchor them. The mentions Flowcraft does earn are being split across what the graph treats as several partially-distinct companies. This is the reason a strong brand can measure weak.TaskHive is the closest comparison the model itself names, and the current category leader. Every row below is a lever Flowcraft can pull. None of them is a product difference.
| Lever | TaskHive | Flowcraft |
|---|---|---|
| G2 reviews | 187 · 4.3★ | 0 |
| G2 category | PM Tools | Workflow Mgmt |
| Gartner Peer Insights | Rated · 41 | Unrated |
| Cited pages naming them (of 24) | 14 | 2 |
| "X vs Y" comparison pages published | 9 | 0 |
| "Alternatives to" pages they appear on | 11 | 0 |
| Wikipedia / Wikidata entity | Both | Neither |
| Profiles linked via schema sameAs | 11 | 0 |
| Case studies with named logos | 23 | 4 |
| Listicle placements in top-5 position | 8 | 0 |
Average day-over-day change in which brands get named, per engine:
ChatGPT carries most of the buyer traffic and moves fast. Claude is slower to move and harder to break into once it settles, which cuts both ways: the hardest engine to enter is the one that holds a position longest once you're in it.
The figures above are noise: the same question answered twice, two days apart. Underneath that noise runs a slower and more consequential movement. A citation that is won is not kept. It ages out of the answer as the pages around it are re-crawled, re-ranked and replaced. Share of the original citation set still being cited, by month:
The red marker is the half-life crossing, not the failure. An alert placed there still has something to re-win; one placed at month 6 is a post-mortem.
Half the set is gone by roughly the ninety-day mark. This is Modeled, not Measured, and the distinction matters here more than anywhere else in the report: Flowcraft has almost no citations to observe decaying, so the curve is drawn from decay observed across other tracked domains, not from Flowcraft's own history. It becomes Measured after one quarter of Flowcraft's own data, which is exactly the point of running it continuously.
Nothing summarised away. Filter by severity, or by the 24 you can close without waiting on anyone. Every gap has an ID, and the 90-day plan below is built out of these IDs rather than a separate list.
Thirty-seven gaps needs an order of attack. These five are where the first ninety days go.
The ranking is the diagonal. Nothing here is ordered by how impressive it sounds; item 4 outranks item 5 because it multiplies the value of every other item rather than adding to it.
Recategorise to PM Tools, rewrite the About text, correct the product taxonomy. Costs an afternoon, and it removes the only actively-wrong information about Flowcraft on a page models read. Everything else on this list is worth less until this is done.
Flowcraft is already listed on the top-ranked page for the head query, just unrated. 20-30 verified reviews moves Flowcraft into the rated tier, which is the filter most cited roundups apply before they list anyone.
Most accept submissions today. Flowcraft Sprints into the six time-tracking-specific lists is entirely untouched ground. The strongest product has zero coverage in its own sub-category.
Pick one canonical name, add Organization schema with sameAs across all 18 profiles, file a Wikidata record. Unglamorous, and it raises the value of every other item on this list rather than adding to it.
"TaskHive vs Flowcraft," "Workstream alternatives." All six competitors already run this play; Flowcraft runs none of it. Pages built with direct quotes, statistics and citations see 30-40% higher generative visibility (Princeton GEO study), so the format matters as much as the topic.
Target: named in at least one of the two shortlist queries by day 90, rated on Gartner PI, and Visibility Index above 30. Steady state is a seat in the five-to-seven-name set, not the top spot, which is a two-year project and a different budget.
This is the one part of the report that isn't measured. Rather than assert a number, here are the inputs. Change them to yours and watch the output move. Every assumption is visible and none of them are ours to make.
best project management software
"Recommended stack: Basecraft/Gridwork + Taskforge + Workstream or Boardwise + ProjectPilot/Chartpath."
is Flowcraft a good option?
"One of the leading horizontal enterprise AI agent platforms." Closest comparison: TaskHive.
About text describes a mattress company. Category: Workflow Management. Reviews: 0. Competitor in same query set (TaskHive): 187 reviews, 4.3★, filed under PM Tools.
Cited-page set collected from live answers, deduplicated by canonical URL, then classified by control class (submission-open / owned profile / editorial / competitor-owned) and checked for the string "Flowcraft" in body text. 2 of 24 matched. Each page was then scored a second time for stance: hostile by name, named but neutral, adversarial without naming, or no engagement, because a name-match check scores a competitor's "alternatives" page as absence when it is the opposite. The decay curve in section 10 is Modeled: it is fitted from citation-set retention across other tracked domains, not from Flowcraft's own history, which does not yet exist.
Run the same prompts again today and the wording will move slightly. That's why every run gets kept, not just the latest. The record is what has value, not the snapshot.
Every gap in the register carries an ID, so a fix shipped in week 3 can be tied to a measurement in week 7, with the day-over-day churn subtracted, so a 9-27% wobble doesn't get claimed as a win.
This runs in the background. No quarterly re-audit required to find out either way, and no arguing about whether the work landed.
Everything above (the 16-phrasing sweep, the 24-page supply audit, the 18-profile check, the crawler and schema pass, the register) runs natively inside Proofsource on a schedule, against your own domain.
Two findings in this report are the argument for that, and neither survives a one-off audit. Half of what phase 1 wins will have decayed before phase 3 closes (section 10), so the fixes need re-verifying on a cycle shorter than their own half-life. And the hostile-citation count is 0 today (section 4), a number whose entire value is being told the week it changes. A document cannot tell you either thing. It was accurate on August 7th and it will go quietly out of date from here.
30 minutes, live in the product. We'll re-run this report in front of you, against today's answers rather than August 7th's.
Book the walkthroughEvery data block in this report carries one of three marks, because "we watched this happen" and "we calculated this from an assumption" deserve different amounts of your trust:
Three live ChatGPT tests, logged out, single run each, Aug 7 2026. A citation-supply audit across 16 phrasings of "project management software," fetching every cited page, deduplicating by canonical URL, and checking each for the vendor names in the tracked set. An automated review-site monitor covering 18 profile surfaces. Volatility figures come from repeated daily sampling of the same prompt set per engine.
What actually decides whether a page gets quoted differs by engine. This is the reference the ranking in Act II is built against:
| Platform | Retrieval & what it cites | Buyer share / volatility |
|---|---|---|
| ChatGPT | Bing-derived index plus its own crawl. Heavy on listicles and review aggregators for "best X" phrasing; quotes the top of a list far more than the tail. | ~60% of assistant-sourced buyer traffic · ~18%/day churn |
| Claude | Reads the Brave index. Slower to change, prefers a smaller and more stable citation set. Harder to enter, stickier once entered. | Smaller share, senior skew · ~9%/day churn |
| Perplexity | Live retrieval on every query, wide citation set including community and video. Most responsive to fresh pages, least stable answer-to-answer. | Research-heavy usage · ~27%/day churn |
| Google AI Overviews | Classic ranking signals still dominate, filtered through the AI layer. Self-published "we're #1" pages still get cited 69% of the time while the answer recommends someone else. | Largest raw reach · moderate volatility |
| Copilot | Bing index plus grounding. Classic Bing SEO signals (IndexNow, Webmaster Tools submission) are the practical lever. | 10-13% of enterprise buyers via M365 · ~16%/day churn |
The reasoning behind "why these levers, in this order" leans on published research, kept separate from anything measured about Flowcraft specifically:
| Source | What it's used for here |
|---|---|
| G2, March 2026 buyer survey | 51% of B2B software buyers now start research in an AI chatbot; 92% say AI shaped their shortlist; 33% bought from a vendor they first discovered in an AI answer. The reason any of this is worth doing. |
| Same survey, on category ownership | Only about 15% of software categories have an entrenched AI-answer "owner" today, but once one forms it holds its lead 90.4% month over month. "project management software" is winnable now; it stays that way for whoever wins it first. |
| Ahrefs, 75,000-brand correlation study | Branded mentions correlate with AI-answer visibility far more than backlinks do (r=0.664 vs. 0.218). Multi-publication syndication lifted citation rates from 7.7% to 34%. Why listicle and press mentions rank above link-building in the plan. |
| Princeton GEO experiment | Pages formatted with direct quotes, statistics and citations saw 30-40% higher generative visibility for lower-ranked sources. Why "how" the comparison pages get written matters, not just that they exist. |
| xfunnel / AI Overviews content-type analysis | Pages that self-publish "we're #1" listicles still get cited by AI Overviews 69% of the time while the same answer recommends a competitor. Why the plan calls for honest comparison pages, not self-ranking ones. |
| llms.txt adoption data | 97% of llms.txt files are never fetched by any AI crawler, and none of the engines above use it as a ranking signal. Explicitly out of scope for this plan. |
| Proofsource platform data | Rolling share-of-voice, answer-engine churn, citation-supply classification and review-site monitoring all come from live sampling, not survey estimates. |
Full evidence ledger, every logged run, kept and available on request. This appendix summarizes it.