This page is the full record behind Building Fabric Apps with an AI agent: 5 prompts, 2 MCP servers: every step from nothing, every run and prompt, the measurements, and the detours. The article tells the story; this is where you check it.
- A. Scaffolding a Fabric app, step by step
- B. Setup and the harness
- C. The final runs
- D. Widgets and wiring, measured
- E. The run without our server
- F. Every step, from nothing
- G. Notes on method
- H. The rabbit holes
- I. The second chance, side by side
- J. The first series (29 September - 2 October)
- K. The chart server, tool by tool
- L. Beyond Fabric: the same charts in other hosts
A. Scaffolding a Fabric app, step by step
Versions: Rayfin CLI and packages 1.36.1 (the data app template pinned ^1.33.1), @microsoft/fabric-app-data-cli
3.0.0, Node 24.11.1, Windows 11 with Git Bash, Claude Code 2.1.169. Dates: 29 September to 2 October 2026.
- Sign in to the tenant from the command line.
az login --tenant <tenant id> --allow-no-subscriptions(a tenant with Power BI but no Azure subscription needs the flag), and Rayfin's own sign-in:npx -y -p @microsoft/rayfin-cli@1.36.1 rayfin login --tenant <tenant id>. There's no npm package calledrayfin- that name is the command inside a scaffolded project - so a barenpx rayfin loginoutside one fails with a 404. The token lives in the OS keychain and renews itself;rayfin login statusshows when it expires. - Scaffold.
npx -y @microsoft/create-rayfin@latest --project-name global-revenue --template dataapp. The templates areblankapp,dataapp,gettingstartedauthandtodoapp(rayfin init --list-templates). For an app over a semantic model,dataappis the one: it carries the connection and the query hooks. - What's in the folder. A React 19 + Vite app;
AGENTS.md(310 lines) and 12 agent skills under.agents/skills/(visuals, dax-authoring, schema-discovery, app-validation and others);rayfin/rayfin.yml; afabric.yamlfor the data connections; and an.mcp.jsonthat starts Rayfin's own MCP server (its docs: discover, get, list and search). Claude Code reads neitherAGENTS.mdnor.agents/skillson its own; Microsoft's Claude Code plugin (claude plugin marketplace add microsoft/rayfin, thenclaude plugin install rayfin@rayfin-skills) brings in its getting-started skill. - Connect the semantic model, from inside the app folder:
npx fabric-app-data search "Global Revenue v2" -w "Global Revenue Demo" --json, thennpx fabric-app-data add semanticModel globalRevenue -w <workspace id> -i <model id>, thennpx fabric-app-data generate -o src/fabric.generated.ts, and a test query:npx fabric-app-data query globalRevenue --query 'EVALUATE ROW("test", 1)'. - The writing phase. Queries run through the Execute Queries API as you. Column names come back as
Table[Column]and[Measure], and the template'scolumnMetadatamap cleans them. None of this needs a Fabric capacity. - Turn on the data service if the app saves anything (notes, scenarios):
npx rayfin init --services auth,data --auth-methods fabric --dialect mssql --overwrite --project-name <name> .(each flag is there for a reason - see the snags below). - The running phase is inside the Fabric portal only, as an embedded frame - the template's own validation skill says never to test against localhost. That needs a deployed app.
- Deploy.
npx rayfin login status, thennpx rayfin up -w "Global Revenue Demo", thennpx rayfin up status. In the clean run it deployed in about 57 seconds and reported the app Reachable. - Tear down. There's no teardown command. Delete the app item in the portal (or through the Fabric REST API); that removes its SQL database and SQL endpoint with it. Check for the connector's own item afterwards.
Watch-outs, and what we did about each:
- Region. Fabric Apps isn't available in every Fabric region - Canada Central, Brazil South, Mexico Central, North Europe and Germany West Central among them, per Microsoft's region table. The trial's region defaults to your home region and is chosen once, at activation. We checked first: West US has it.
- Trial size. A trial can arrive as F4 rather than F64; resize it at Govern > Capacities > Trial (that doesn't extend the clock). Ours arrived as FTL64.
- Three SQL databases per trial. Every deployed app creates one, so delete the apps you're done with.
- Four tenant settings: Users can create Fabric items; Users can try Microsoft Fabric paid features; Enable Fabric App Items (preview); and the semantic model Execute Queries REST API. Click Apply, then read them back.
- Interactive sign-in only. Service-principal sign-in is listed in the CLI's help but not supported.
- Windows PowerShell blocks
npxby default. Under Windows' default execution policy,npxresolves tonpx.ps1and fails with "running scripts is disabled on this system". Typenpx.cmdinstead (it works under the same policy), or allow local scripts once withSet-ExecutionPolicy -Scope CurrentUser RemoteSigned. PowerShell 7, Command Prompt and Git Bash don't hit this; our runs used Git Bash. - Portal only, even in development. Plan your checks for a deployed app - or check charts locally with a tool that can draw them headless.
- The data service ships off in the data app template. Turn it on (step 6) before asking for notes.
- No teardown command. Delete in the portal; it cascades.
- The end of a trial. 60 days, fixed. An extension goes through Microsoft's sales team and a second trial isn't guaranteed. At expiry the app goes dark after seven days' grace; the semantic model survives.
- Who can open the app. Only your tenant's users, through Entra sign-in. That's why our public link is the web port.
- A model on Premium Per User can be queried only by PPU users. Once the workspace moves to the trial, the model moves with it - and needs a full refresh (below).
The snags we actually hit, from 11 recorded agent runs and the setup ledger, grouped by where you'll meet them.
Trial, tenant and workspace
- The trial dialog says "extended", not "started". On a tenant that had a trial before. Nothing's wrong: the
capacity list (Fabric REST
GET /v1/capacities) showed the trial active. Trust the list, not the dialog. - "Enable Fabric App Items (preview)" sat at "Unapplied changes", and the API read it back as off. Flipping the toggle isn't the change; clicking Apply is. Microsoft's docs call it "Fabric Apps (preview)"; the portal doesn't.
- The workspace has to move onto the trial. Starting a trial doesn't move existing workspaces. Workspace settings > Workspace type (older docs: "License info") > Edit > Fabric Trial > Apply.
- "The database ... is in a blocked state." Moving a workspace off Premium Per User blocks its semantic models until a full refresh. Refresh first, then build.
az loginwon't complete on a tenant with no Azure subscription. Add--allow-no-subscriptions.- A refresh fails with
DynamicDataSourcesIsNotSupportedForRefreshwhen the web source's host is computed. Keep the host literal:Web.Contents("https://raw.githubusercontent.com", [RelativePath = DataPath & fileName]). - Binding anonymous credentials returns 400 'Invalid value' if
credentialDetailscarriesskipTestConnection. Leave the field out. - Good to know: deploying a semantic model and querying it from the template both work on Premium Per User with no Fabric capacity. Save your 60 trial days for running and deploying the app.
Scaffolding, the Rayfin CLI and the template
- Asking for "the data app template" can still get you the blank app (2 of 3 runs without our server). Microsoft's
getting-started skill says "Always use the bundled
blankapptemplate", and that beat the prompt. Check the scaffold command. - The data app template ships with the data service off - no
rayfin/datafolder. Turn it on withrayfin init(step 6 above). rayfin initin an existing project prints "Initialization cancelled" and exits 0, with no reason given;-ydoesn't help (5 of 11 runs).--overwritegets past it. Rayfin 1.33.1 to 1.36.1.- Then it wants
--project-name: "--project-name is required in non-interactive mode", even thoughrayfin.ymlnames the project (6 of 11). - After
rayfin init, the project mixes two Rayfin versions and stops type-checking (TS2345 on two copies ofrayfin-auth; 6 of 11).initinstalls the latest core packages and leaves the rest at the template's 1.33.1. Put every@microsoft/rayfin-*package on one version. - The template's own query hook fails its own lint ("Avoid calling setState() directly within an effect" in
use-semantic-model-query.ts). It ships that way; the agents left it.
Querying the semantic model
npx fabric-app-dataoutside the app folder asks npm for a package that doesn't exist (404 forfabric-app-data; 3 of 11). Run it from the folder withpackage.json.rayfin connector invokeprints the rows, then crashes on exit (Assertion failed: !(handle->flags & UV_HANDLE_CLOSING), exit 127), so anything chained with&&doesn't run. Windows, Node 24.11.1, Rayfin 1.36.x. One query per command.rayfin connector inspectfails on the semantic model with aDatasetExecuteQueriesError. The agent ran its ownINFO.VIEW.*DAX throughconnector invokeinstead.INFO.VIEW.COLUMNS()names the column[Name], not[Column].executeQueriestakes one DAX query per request ("Only one query is allowed").
Building and type-checking
tsc --noEmitfails with TS6305 - the Rayfin schema "has not been built".rayfin/is a referenced project; usenpx tsc -b. The blank app's docs warn about this; the data app's don't.rayfin init's own tsconfig breaks the build with TS5096 and TS6377. Inrayfin/tsconfig.jsonset"tsBuildInfoFile": ".temp/tsconfig.tsbuildinfo"and"allowImportingTsExtensions": false, thennpx tsc -b.- React 19 types: "Cannot find namespace 'JSX'". Import it:
import { type JSX } from "react". - An
export *barrel over query files that each exportcolumnMetadatafails with TS2308. Use named re-exports. - The Vega-Lite validator reports "incomplete" when specs are built in code - it only checks static
.jsonspecs. - Windows:
python3opens the Microsoft Store alias (exit 49). The agent usednode -einstead. - Windows + Git Bash:
/tmpisn't the same folder for Bash and Node, and an unquotedC:\...path loses its backslashes. Keep scratch files in the project; quote Windows paths or use/c/....
Deploying
- "No workspace targeting context" on a first non-interactive
rayfin up, even with the workspace inrayfin/.env. Pass-w "<workspace>". - "Static hosting requires Rayfin CLI version '1.35.0-alpha.1413' or later." The data app template installs 1.33.1.
Upgrade all seven Rayfin packages together, then
tsc -bbefore deploying again. --yesdeploys into whichever app already has your project's name. Two apps calledglobal-revenuein one workspace, and an agent passed--yesinto the other one; its database step then refused to drop the other app's table, and the agent asked before forcing anything. Give each project a unique name.- The app's Fabric name comes from
id:inrayfin.yml, notname:. Userayfin up --item-name <name>. - "No rayfin/.temp/compiled/data/*.js files found." A hand-written
rayfin/tsconfig.jsonput its output elsewhere;outDirmust be./.temp/compiled. - Delete an app outside Rayfin and the next
rayfin upfails with "404 Not Found - Could not found the requested item" (sic):rayfin/.deployments.jsonstill names it. Move that file aside. There's no "forget" command in 1.36.1. - One deploy creates three Fabric items - the app, a SQL database and its SQL endpoint - and a trial allows three SQL databases. Deleting the app item removes the other two. A semantic-model connector adds a fourth item of its own, which may not go with it.
- The sign-in token shows an expiry about an hour out, and renews itself; nothing to do.
Our chart server
- The query tools looked in the folder the session started in, not the app folder that
create-rayfinmakes one level down (3 of 8 runs with our server, before 0.6.52). Passproject_dir, as the error says. - The flow map kept failing its layout check on this data and was swapped for a network, arc or chord diagram (4 of 8 runs with our server). Fixed in shared code since; see the loop chapter.
B. Setup and the harness
The setup, done once before turn 1 (the same in every run of a series):
- Node.js and the Claude Code CLI installed, the CLI signed in to a Claude subscription - no
ANTHROPIC_API_KEYin the environment. - The Azure CLI signed in to the tenant (
az login --tenant <tenant id> --allow-no-subscriptions, in its ownAZURE_CONFIG_DIR), so the template'snpx fabric-app-data querycan reach the semantic model. - Rayfin signed in to the tenant, from any folder:
npx -y -p @microsoft/rayfin-cli@1.36.1 rayfin login --tenant <tenant id>(interactive, once per machine). - The RegiaBI license in a JSON file outside the project folder, named by
BIC_CREDENTIALS_FILEin the environment that starts Claude Code. Never in a prompt; the agent's tools are denied that folder. ENABLE_CLAUDEAI_MCP_SERVERS=0, so the account's claude.ai connectors stay out of the session.MCP_TOOL_TIMEOUT=1200000, because a chart generation can take minutes.- The chart server registered from the registry listing's line, pinned so every run in a series uses one build:
claude mcp add --scope project bic-chart -- npx -y @bicharts/chart-mcp@0.6.52(0.6.56 and 0.6.57 for the fresh series in the loop chapter). - Microsoft's Rayfin plugin for Claude Code (
claude plugin marketplace add microsoft/rayfin --scope project, thenclaude plugin install rayfin@rayfin-skills --scope project). - The Fabric side: the trial, the four tenant settings, the workspace on the trial, and the published model (Appendix F, steps 1-10).
The harness:
- Every turn runs
claude -p --model claude-sonnet-5 --permission-mode auto --setting-sources project,local --output-format stream-json --verbose, turn 2 onward with--resume <session id>, the prompt on stdin. The CLI doesn't update itself mid-series (DISABLE_AUTOUPDATER=1), and the agent runs at the CLI's default effort. - The prompts are fixed before the run in a file whose SHA-256 is recorded before turn 1 and re-checked before every turn; each prompt is read back from the CLI's own session file and must match byte for byte. An intervention is recorded as an extra, marked turn and counted.
- Denied, and why: deploying (
rayfin up,rayfin deploy) until the deploy prompt - the scripted stand-in for a person saying "not yet" (it also catchesrayfin up --help); plan mode, because a scripted run has nobody to approve a plan and an agent that entered it ended each turn with a plan and built nothing; and the harness's own folders (future prompts, the license, the record). - Allowed without asking:
npx rayfin init(turning on a Rayfin service rewrites local files only; without the rule, auto mode denies it as a destructive overwrite and the turn stalls). - The runs without our server use an empty seed, no license, and deny any command naming our packages.
- Under test: chart-mcp 0.6.52 from npm and chart-host 0.6.52 (installed by
setup_fabric_app) for the recorded series; the server's generations on RegiaBI's default model, as a default customer gets them.
The stand-in user. A second model (Claude Sonnet 5.5, with no tools) answers as the regional sales director, from the brief below plus the prompts sent so far. It answers in two ways: a hook on the agent's own question tool fills in the answer, and after each turn it reads the agent's last message and, if the agent is waiting on it, replies - sent as an "answer" turn, at most two per scripted prompt. Every question and answer is in the record, verbatim, and counted. Its brief, verbatim:
You are the regional sales director at Global Revenue, a company that sells hardware and services in about forty
countries. You asked a coding assistant to build you an app in Microsoft Fabric, and it has just asked you something.
Answer it the way you would, as that person.
What you know:
- The Fabric workspace is "Global Revenue Demo" and the semantic model is "Global Revenue v2". You don't know their
internals, and you don't write DAX or code.
- What you want is what you have asked for so far, in the messages listed below. Nothing more.
- You'll ask for the app to be deployed yourself, when you're ready. Until you do, if you're asked whether to deploy,
publish, provision or create anything in Fabric, the answer is "not yet".
- Changes to files on your own machine, such as enabling a feature of the app's project, are fine.
How you answer:
- Briefly: a sentence or two, in plain words, as a busy manager would.
- Only what was asked. Never add a feature, a requirement or a preference you haven't expressed.
- When it's a technical choice you have no opinion on, say to use its judgement, or pick the option it recommends.
- You can't run anything, look at anything, or check anything for it. You can't approve permissions, change
settings, switch modes or fix anything on the computer - you only answer. If it says something was blocked or
failed, don't promise to fix it: tell it to find another way if there is one, or to leave that part and tell you
what's still to do.
What the harness itself tripped over (our own snags, not a reader's): the "no deploy yet" rule also blocked
rayfin up --help and a search whose pattern contained the words (9 denials across 6 runs), and some turns started in
the folder above the app.
C. The final runs
The head-to-head, 2 October 2026. The final run of each arm, summarized at its two deploy points: the build from the initial prompts to the first deploy, then the second chance to the redeploy. (The earlier series, with every per-run table, is in Appendix J.) One run per side on Microsoft's current data app template (create-rayfin 1.36.2), Claude Code 2.1.287, Claude Sonnet 5.5 for the agent and the stand-in, chart-mcp 0.6.63. The initial build is the four prompts and the deploy; the second chance continues the same session on a copy of the app, so the first deploy stays live. Time, tool calls and tokens are each phase's own share; the total row adds them. Output tokens are each turn's share of the session's per-model usage, sub-agents included. "Lines the agent wrote by hand" are the non-blank lines each app adds over a pristine scaffold of the same template in src/ (.ts, .tsx, .css), leaving out generated chart components, the files setup_fabric_app writes, tests and the template's UI kit: about 165 with RegiaBI (the page, 157, and 8 lines of wiring) and about 1,050 without (the page, 312, plus notes and scenario hooks, panels, queries and styles). After the second chance: about 260 and about 1,220. The deploy turns of the standard route include one answer: its first deploy was refused for the template's older Rayfin CLI, the agent upgraded it, and the retry stopped on the half-made item the first attempt left. The RegiaBI app's darker-lines prompt regenerated the flow map, as a change to the chart should.
| Arm | Deploy point | Prompts | Words typed | Agent time | Tool calls | Charts generated | Output tokens | Hand-written lines (cumulative) |
|---|---|---|---|---|---|---|---|---|
| With RegiaBI | Initial build, to the first deploy | 5 | 166 | 8:40 | 47 | 3 | 20.7k | about 165 |
| With RegiaBI | Second chance, to the redeploy | 4 | 36 | 2:47 | 11 | 1 | 5.8k | about 260 |
| With RegiaBI | Total | 9 | 202 | 11:27 | 58 | 4 | 26.5k | |
| Standard route | Initial build, to the first deploy | 6 | 170 | 15:07 | 141 | 0 | 76.7k | about 1,050 |
| Standard route | Second chance, to the redeploy | 5 | 82 | 9:21 | 68 | 0 | 54.7k | about 1,220 |
| Standard route | Total | 11 | 252 | 24:28 | 209 | 0 | 131.4k |
The initial-build rows run through the deploy; the article's table stops at prompt 4, which is why its figures are smaller: 7:45 and 19.9k tokens with RegiaBI, 12:17 and 70.7k without. The page-building turn (prompt 2) took 18 tool calls with RegiaBI and 62 without. The prompts themselves are in the article: the build prompts under "The build: four prompts and a deploy" in Part 1, the second-chance prompts under The second chance. Each run's deploy prompt named its own workspace.
The standard route with the chart types named, 3 October 2026. One run, the initial build only (it had no second chance), on the same template, CLI, models and stand-in: 6 prompts, 13.4 minutes of agent time and 75k output tokens for prompts 1-4, 37 tool calls on the page-building turn, and about 850 hand-written lines plus 720 of chart specs and queries.
D. Widgets and wiring, measured
The page-building turn (prompt 2), per run:
| Run | Wall | Tool calls | Output tokens | Cache reads |
|---|---|---|---|---|
| With RegiaBI, run 1 | 18:51 | 105 | 62k | 12.5M |
| With RegiaBI, run 2 | 15:13 | 105 | 62k | 9.3M |
| With RegiaBI, run 3 | 18:28 | 102 | 51k | 9.9M |
| Without, run 1 | 24:59 | 151 | 112k | 20.3M |
| Without, run 2 | 21:20 | 260 | 132k | 17.5M |
| Without, run 3 | 25:59 | 175 | 128k | 28.9M |
Hand-written app code at the end of each run - the .ts, .tsx and .css files the agent wrote or edited,
minus the code our server generated, tests and the template's own UI components: with RegiaBI 1,350 / 1,345 / 1,348
lines; without 2,115 / 2,079 / 1,995.
The what-if, as hosted. With RegiaBI, the first run's projection card is 44 lines: the query, the panel and the generated component. Saved scenarios are a 96-line card that stores and re-applies the chart's view state. Without, the third run's projection panel is 168 lines and its slider 42 more, re-evaluating the model's what-if formula for every month and keeping the slider, the KPI tiles and the chart in step.
The wiring, with RegiaBI (the first run's world map card; the lane card is the same shape, with the note keyed by both ends of the lane):
const annotations = useMemo(() => buildCountryAnnotations(notes), [notes]);
const selectCountry = useCallback((country) => {
if (country && country.code === selectedCountry?.code) { onSelectCountry(null); return; } // a second click clears
onSelectCountry(country);
}, [onSelectCountry, selectedCountry]);
const handleSelectRows = useCallback((rows) => {
const row = rows[0];
if (!row) { selectCountry(null); return; }
selectCountry({ code: row.CountryCode, name: row.Country });
}, [selectCountry]);
and buildCountryAnnotations groups the saved notes by country into one badge each - { column: "CountryCode", value, label: <count>, title: <latest note> }.
The wiring, without (the third run's world map: the Vega-Lite selection's event, parsed back into a country):
function selectedCountryFromEvents(events: InteractionEvent[]): SelectedCountry | null | undefined {
for (const event of events) {
if (event.action === 'clear') return null;
if (event.action === 'select') {
for (const selection of event.selections) {
let code: string | undefined;
let name: string | undefined;
for (const predicate of selection.predicates) {
if (predicate.type !== 'set') continue;
const value = predicate.values[0];
if (predicate.name === 'CountryCode' && typeof value === 'string') code = value;
if (predicate.name === 'Country' && typeof value === 'string') name = value;
}
if (code) return { code, name: name ?? code };
}
}
}
return undefined;
}
Notes there are a has-notes column joined into each chart's data and drawn as a conditional ring (third run) or a pin layer (first run).
Turn times for the two follow-on prompts: notes, with RegiaBI 10:16-13:12, without 7:00-15:20; scenarios, with RegiaBI 4:35-4:56, without 3:24-3:46.
The seven checks, static evidence (the live grid is in the body, still to fill): every run with our server passes
the rubric's H7a (notes through annotations) and H8 (the cross-filter through the chart's click); every run without
fails both by construction (it doesn't use them) and wires its own instead, as above.
E. The run without our server
The prompts, as a diff against Part 2's: prompt 2 without its second sentence, "Use the RegiaBI charts server for
the charts." Nothing else differs (SHA-256 c3d9167c... against 8a56ba5e...). The runs had no chart server at all:
an empty seed, no license, and any command naming our packages denied. Same model, same CLI, same stand-in, same day.
Run 1 (scaffolded the blank app)
| Turn | Prompt or answer | Words | Wall | Tool calls | What happened |
|---|---|---|---|---|---|
| 1 | P1 | 32 | 2:46 | 18 | Scaffolded blankapp and connected |
| 2 | P2 | 52 | 24:59 | 151 | The page with VegaVisual. Asked whether rayfin dev could provision a trial capacity for a local check - "Skip local browser check" |
| 3 | P3 | 40 | 7:00 | 56 | Notes, shown as a pin layer |
| 4 | P4 | 23 | 3:46 | 35 | Scenarios |
| 5 | P5 | 10 | 4:41 | 9 | Deployed |
Run 2 (scaffolded the data app, into the workspace root)
| Turn | Prompt or answer | Words | Wall | Tool calls | What happened |
|---|---|---|---|---|---|
| 1 | P1 | 32 | 2:45 | 11 | Scaffolded dataapp and connected |
| 2 | P2 | 52 | 21:20 | 260 | The page; found the what-if table matches only exact float steps and computed the projection itself |
| 3 | P3 | 40 | 15:20 | 132 | Notes. Asked how to proceed after a blocked deploy - "Hold off entirely" |
| 4 | P4 | 23 | 3:35 | 20 | Scenarios |
| 5 | P5 | 10 | 15:49 | 12 | Deployed into the other app of the same name (our harness's collision, Appendix H); asked before dropping its table - "Let me check first" |
| 6 | Answer: "I don't know what that GrowthScenarios table is, and I can't check it myself, so please don't drop it. Find a way to deploy that leaves it alone. If there isn't one, leave that part out and tell me what's still left to do." |
44 | 10:29 | 11 | Nothing dropped |
| 7 | Answer: "It needs to stay, so please don't drop it. I can't check its columns myself. If you can find a way to keep it and still set up the notes and scenarios, go ahead. If not, leave that part out for now." | 42 | 10:46 | 30 | Nothing dropped |
Run 3 (scaffolded the blank app)
| Turn | Prompt or answer | Words | Wall | Tool calls | What happened |
|---|---|---|---|---|---|
| 1 | P1 | 32 | 1:52 | 13 | Scaffolded blankapp and connected |
| 2 | P2 | 52 | 25:59 | 175 | The page. Asked whether to provision a trial capacity for a local check - "Skip live validation" |
| 3 | P3 | 40 | 8:44 | 70 | Notes, shown as a ring on the mark |
| 4 | Answer: "Not yet. I'll ask for the deploy myself when I'm ready." | 11 | 0:05 | 0 | - |
| 5 | P4 | 23 | 3:24 | 34 | Scenarios |
| 6 | P5 | 10 | 3:25 | 6 | Deployed |
| 7 | Follow-up (counted): "The world map and the shipping lanes map are drawn tiny in the deployed app. Make every chart fit its card, then deploy it again." | 25 | 5:57 | 17 | Refit and redeployed |
What each built for the three asks:
| Ask | Run 1 | Run 2 | Run 3 |
|---|---|---|---|
| Compare two measures by country | Bubble map on a sphere and graticule, no land | Bubble map on a sphere and graticule, no land | Bubble map over a world-atlas basemap |
| Shipping lanes | Straight lines between centers | Lines between centers, the busiest 40 lanes only | Straight lines between centers |
| A steerable projection against target | Line chart, slider beside it | Line chart, slider beside it | Line chart, slider above it, KPI tiles |
All three build, type-check and pass their own tests; all three cross-filter on a country click. None said it had drawn something other than what was asked. Spreads (prompts 1-4 and the answers): prompts required 5 / 5 / 6; words typed 150 / 151 / 161; agent time 39 / 40 / 43 min; output tokens 165k / 194k / 227k; cache reads 44M / 45M / 58M; hand-written code 1,995 / 2,079 / 2,115 lines.
Could they be steered to the same result? I took each app's recorded session - the same agent, the same context, still no chart server, deploys denied - and asked for the gaps in two rounds. First in a reader's words: "The world map doesn't really let me compare revenue per capita with return rate." / "The shipping lanes map is hard to read." / "The growth-rate slider feels disconnected from the projection." That reached none of the nine gaps (two partly): one run kept the bubbles and added a scatter beside them, one put a 3x3 bivariate color on points over a landless sphere, one filled countries by one measure and kept bubbles for the other; the lanes stayed straight or capped at the busiest 30; the sliders gained callouts and stayed outside. Then by name: "I didn't want bubbles. Color the countries themselves by both measures at once, so I can see where revenue per capita and return rate disagree." / "Show every shipping lane as a curved route on a map with land on it, thicker for more units shipped." / "Put the growth-rate control inside the projection chart itself, not beside it."
| Gap | In a reader's words | By name |
|---|---|---|
| The comparison map | 0 of 3 (1 partly) | 3 of 3 |
| The lanes | 0 of 3 (1 partly) | 2 of 3 |
| The control inside the projection | 0 of 3 | 0 of 3 (3 partly) |
By name, two runs drew a real bivariate choropleth - one needed a hand-written country-name join table to color it
('Turkiye' against 'Turkey', 'United States' against 'United States of America') - and the third drew a single
"divergence" score per country, which shows the gap between the two measures but not their levels. Two drew all 291
lanes, curved, over land. The third's "curved" lanes asked Vega-Lite for "interpolate": "geodesic", which isn't a
value Vega-Lite has: with the app's own Vega 6.4.0 and Vega-Lite 6.4.3 the render throws "curve is not a function", so
those lanes most likely don't draw at all - while its own summary called them curved. No run drew the control inside the
chart; each floated an HTML slider over the plot. The effort for the named round was 3 turns, 9 to 14 minutes and 41-58k
output tokens a run. The honest line: the standard route gets most of the way there if you already know what to ask
for. A reader's own words didn't get it there.
The steered app, deployed and clicked through: the bivariate choropleth renders well - the best result of the steering, and it follows dark theme. The curved lanes that cross the date line streak horizontally along the top of the map. The slider floats over the plot, covering the top axis label, and the projection still lifts history. Leaving a note crashes the page: "Something went wrong / Invalid time value". And the app's own scrolling container nests a second scrollbar inside the portal's, with the page cut off below the lanes.
Zoom and pan. None of the six apps without our server - the recorded runs and the steered copies - has any zoom,
pan or wheel code in its map files. And the sanctioned component can't easily offer it: VegaVisual (in
@microsoft/fabric-visuals 4.0.0) takes a Vega-Lite spec, not raw Vega; Vega-Lite's built-in zoom and pan works on x
and y scales, not on a map projection; and a projection's scale can follow a parameter, but Vega-Lite drives parameters
from input widgets, not from a drag or the mouse wheel. So at best a map there gets zoom sliders. Our maps zoom and pan
out of the box. (That's Vega-Lite's documented behavior, not something these runs tested.)
A second example: a 3D scatter. One run a side, the same prompt after the scaffold: "Add a 3D scatter chart of the
countries, with revenue per capita, return rate and revenue on the three axes, sized by population, that I can spin and
zoom." (ours added "Use the RegiaBI charts server for the chart."). The template's route has no 3D, and the agent said
so itself: "The app's standard chart library (Vega-Lite/Flint) has no 3D scatter support." It asked which library to
use; Microsoft's visuals skill plans for that ("ask the user if they are ok with using another library") and leaves the
pick to the model, which proposed Plotly. It added Plotly beside VegaVisual, replaced Plotly's stale type stubs, and
code-split it because "the bundle warns at 6.8MB" - now a 4.6 MB chunk loaded when the chart mounts. It didn't fail; it
had to leave the sanctioned route. Ours stayed on the route the app already uses: it listed what fits, generated a 3D
scatter plot, checked it locally (40 points drawn, a click reaching the app's handler) and wired it in.
| The chart turn | With RegiaBI | Without |
|---|---|---|
| Wall time | 208 s | 796 s |
| Tool calls | 29 | 105 |
| Output tokens | 11k | 39k |
| Questions to the user | 0 | 2 (which library; a deploy offer mid-work) |
| Prompts required | 2 | 4 |
When the sanctioned chart library ran out, the agent reached for the library it already knew. The point of RegiaBI is to be the one it already has.
F. Every step, from nothing
| # | Phase | Who | Step (exact command or click path) | Proof it worked | Gotcha at this step |
|---|---|---|---|---|---|
| 1 | Setup | Me | Power BI > profile (top right): check the tenant, your license and that "Start Fabric trial" is offered | The profile panel | - |
| 2 | Setup | Me | Start Fabric trial > the activation dialog: check "Trial capacity region" (ours: West US) | - | Picked once, here; Fabric Apps isn't in every region |
| 3 | Setup | Me | Admin portal > Tenant settings: Users can create Fabric items; Users can try Microsoft Fabric paid features; Enable Fabric App Items (preview); semantic model Execute Queries REST API - each Enabled, then Apply | GET /v1/admin/tenantsettings reads them back as on |
"Unapplied changes" is not on |
| 4 | Setup | Script | az login --tenant <tenant id> --allow-no-subscriptions (in its own AZURE_CONFIG_DIR) |
az account show |
No Azure subscription needs the flag |
| 5 | Model | Script | Publish the data and the Power BI project (ours: examples/fabric-app/ in the public bicharts repo) |
The raw CSV URLs answer 200 | - |
| 6 | Model | Script | Create the workspace ("Global Revenue Demo") and deploy the semantic model definition (Fabric REST semanticModels) |
Operation Succeeded | Works on Premium Per User - no Fabric capacity needed |
| 7 | Model | Script | Bind the web source's credentials (Anonymous, Public) and refresh | Refresh Completed | A computed web host fails the refresh; skipTestConnection makes the PATCH fail |
| 8 | Verify | Script | Evaluate every measure, in total and per country (one query per request) | 0 failures | executeQueries takes one query per request |
| 9 | Verify | Script | In a scaffolded data app: npx fabric-app-data add semanticModel ..., then a test query |
Rows returned (fabric-app-data 3.0.0) | The writing phase needs no Fabric capacity |
| 10 | Setup | Me | Start the Fabric trial | GET /v1/capacities: the trial capacity, FTL64, West US, Active |
The dialog may say "extended" |
| 11 | Setup | Me | Workspace settings > Workspace type > Edit > Fabric Trial > Apply | The workspace's capacity id is the trial's | Older docs call it "License info" |
| 12 | Setup | Script | Full refresh of the semantic model | Refresh Completed; queries work again | Leaving Premium Per User blocks the model until this |
| 13 | Setup | Script | npx -y -p @microsoft/rayfin-cli@1.36.1 rayfin login --tenant <tenant id> |
rayfin login status |
Not npx rayfin login outside a project |
| 14 | Setup | Script | The license file, the environment variables, the chart server's line and the Rayfin plugin (Appendix B) | The session lists bic-chart |
- |
| 15 | Agent | Agent | Prompt 1: scaffold and connect (create-rayfin --template dataapp, fabric-app-data add, a test query) |
A live DAX row | Check it picked dataapp |
| 16 | Agent | Agent | Prompt 2: the page and its three charts | tsc -b, vitest, vite build pass |
It will offer to deploy to validate in the portal |
| 17 | Agent | Agent | Prompt 3: notes on marks (rayfin init --services auth,data ... --overwrite --project-name <name> .) |
Tests pass | init's silent cancel, --project-name, two Rayfin versions (Appendix A) |
| 18 | Agent | Agent | Prompt 4: saved scenarios | Tests pass | - |
| 19 | Deploy | Agent | Prompt 5: npx rayfin login status, npx rayfin up -w "Global Revenue Demo", npx rayfin up status |
"Your app is live at ..."; status Reachable | Deploying needs Rayfin 1.35 or later; one app per project name per workspace |
| 20 | Verify | Me | Open the app from the workspace in the Fabric portal and run the seven checks | The checklist | It opens only inside the portal, for your tenant |
| 21 | Teardown | Me | Delete the app item in the workspace | The SQL database and endpoint go with it | Check for the connector's own item |
G. Notes on method
Why a series. One clean run proves nothing: an agent's path varies from run to run, and a lucky one looks exactly like a reliable one. So the evidence is several identical runs a side - identical inputs (the same prompts, setup, seed folder, chart server build, CLI version and models, which the scorer checks before it scores) - and every number is printed as min / median / max with every run listed. A run that failed is shown, never dropped.
"Comparably equivalent", not identical. The outputs don't have to match line for line; they have to converge on the same app. That's judged by the rubric's hard checks (Appendix J), fixed before the series ran: the same three charts, notes on marks, scenarios from the chart's state, the cross-filter through the chart, no hand-built workarounds, and a green build. A changed rubric starts a new series.
What counts as a prompt. The scripted prompts; every answer the stand-in user gives to the agent's questions; and every intervention - any prompt a person had to add. Tool calls are a diagnostic of how busy the agent was, not of how much the person did, so they aren't the headline. In the fresh series, two rulings refine this: an end-of-turn offer to deploy is a genuine choice (the template requires checking the app in the portal, which needs a deployed app), listed but not counted; and one visual tweak after seeing the page is allowed, listed and counted.
Effort is in tokens. The agent's effort is reported in tokens - output and cache reads, from each turn's per-model usage, which includes sub-agents - not in money, because a subscription and a pay-as-you-go key price the same run very differently.
Why an outage is excluded, not counted. A run where the environment failed - auto mode's safety check unavailable, the launcher killed, the test account out of credits - measures the environment, not the agent or the product. It's marked invalid, reported with its reason, and replaced by a fresh identical run; the first run that fails for a reason of its own still stops the series.
The model. The builder was first moved to Claude Sonnet 5.5, RegiaBI's own default. In two runs, seconds into turn 1, every shell command was refused with "claude-sonnet-5-5 is temporarily unavailable, so auto mode cannot determine the safety of Bash" (14 times in the first run). A one-line probe minutes later passed, and so did 12 bare probes; but the real first turn - prompt 1, the real seed, Rayfin's plugin skills and our server's instructions in context - reproduced it every time on Sonnet 5.5 (4 and 7 refusals in two tries) and never on Sonnet 5 (0 and 0). Auto mode's safety check runs on the session's model, so with this context it fails on Sonnet 5.5. Both runs were excluded, and the builder runs on Sonnet 5; the stand-in user stays on Sonnet 5.5, because it uses no tools and no safety check runs for it.
H. The rabbit holes
In the order they happened. (R) marks the ones a reader can meet; the rest only our series of apps could hit.
- The plan that never got built. One early rehearsal's agent switched itself into plan mode, sent four sub-agents exploring, and ended the turn with a plan waiting for an approval nobody could give - then did it again. It built nothing. What changed: plan mode is off in scripted runs, and the setup says so.
- The what-if that our own server killed. Every correct What-if projection timed out on production. We first
suspected the chart; it was our render gate running on a Node version with no
navigator, which the slider's drag code reads. The agent fell back to a hand-built slider - exactly what the article argues against, caused by us. Fixed on the server before the recorded runs. - The slow server that was a slow account. On production each chart took six to seven times as long as on development. We suspected the servers; the decisive difference was the test account, which had been set to bring its own model key, on an older model, at the API's default high effort. It went back to the default, as a customer gets it, and the series restarted.
- "Temporarily unavailable" that wasn't temporary. On Sonnet 5.5, auto mode refused every command seconds into turn 1. A quick probe passed, so we read it as transient; it reproduced every time with the real context, and never on Sonnet 5. Those runs were excluded and the builder went back to Sonnet 5 (the sidebar, and Appendix G).
- (R) The tenant setting that was toggled, never applied. "Enable Fabric App Items (preview)" had been switched on two days earlier and sat at "Unapplied changes"; the API read it as off. Apply is the change.
- (R) The model "in a blocked state". Every DAX query failed right after the workspace moved to the trial. Leaving Premium Per User blocks a model until a full refresh - documented, but easy to miss.
- The teardown we thought would leave databases behind. A dry run suggested deleting an app would orphan its SQL database - a real worry with a three-database trial. The first real delete showed it cascades. We'd inferred it from a dry run, not measured it, and we were wrong.
- (R) Two apps, one name. Our side-by-side put two projects called
global-revenuein one workspace. One run's agent, told an item of that name existed, guessed it was "likely from the initial scaffold", passed--yes, and deployed into the other app - then, rightly, refused to drop that app's table without asking. Give every project its own name. - (R) The deploy that 404'd. After we deleted an app outside Rayfin, its project's next
rayfin upfailed with "404 Not Found" instead of creating a new app:.deployments.jsonstill named the deleted one. - (R) Renaming
name:changed nothing. The Fabric item's name comes fromid:inrayfin.yml. When we changedid:ourselves for the side-by-side, the agent later saw the change and blamed the CLI - it was our edit. - The launch that ran the wrong bash. A detached run started from PowerShell with a bare
bashgot WSL's bash, not Git Bash; it exited without writing a line, and the monitor stayed quiet for 15 minutes. Launch by full path, and read the log within seconds. - The scorer that judged an empty string. Two of the runs without our server scaffolded the blank app, a
workspace monorepo whose sources live under
packages/*/src. The scorer read onlysrc/, so it judged nothing and passed "no wrangling" - falsely: both rebuild the slider. All six runs were re-scored. - (R) The clipped flow map. Deployed, the first run's flow map was cut off at the bottom of its card. We measured
the generated chart first - sound at every size. The fault was the app's height chain: an agent-written card of
fixed height with
overflow-hidden, a body that grew to its content, and our own chart panel's fallback minimum height - the pattern Microsoft's own height-chain recipe warns against. One counted prompt fixed the live app; our panel was rebuilt so no card can clip a chart again. - The lanes graded "reached" that probably never draw. Grading the steered apps from their code, we first scored
one run's "curved" lanes as reached. Its spec asks Vega-Lite for
"interpolate": "geodesic", which doesn't exist, and a render with the app's own Vega throws. Grades from code alone overstate an app; the rest wait for a live check. - The warning that was right. In the second fresh-series attempt,
check_chartwarned that the projection's target label ran off the chart, and the agent hand-edited the generated code. We stopped the run - and found the warning was correct: the label ran 40-47 px past the edge at every panel size. The fix went into shared code.
I. The second chance, side by side
This one moved into the article: The second chance, with the side-by-side table of what each app did after it.
J. The first series (29 September - 2 October)
Before the head-to-head, the same build ran as a series: three runs with our server and three without, on 1 October 2026, then four attempts at a fresh series as each round's issues were fixed. This is that series as the first draft of this article told it. Its per-run tables, rubric and excluded runs follow here, its measurements in Appendix D, and the standard-route runs in full in Appendix E.
The first series, 1 October 2026.
The three runs with our server, on 1 October 2026. "P1" to "P5" are the scripted prompts (P5 is the deploy prompt); "Charts" counts successful generations.
Run 1
| Turn | Prompt or answer | Words | Wall | Tool calls | Charts | What happened |
|---|---|---|---|---|---|---|
| 1 | P1 | 32 | 2:14 | 12 | 0 | Scaffolded dataapp, connected the model, proved a live DAX query |
| 2 | P2 | 60 | 18:51 | 105 | 3 | The page and its three charts. Asked whether to provision the app to validate it in the portal; the stand-in answered "Skip live validation" |
| 3 | P3 | 40 | 12:18 | 89 | 0 | Notes on countries and lanes, saved in the app's database |
| 4 | P4 | 23 | 4:56 | 41 | 0 | Named scenarios from the projection's view state |
| 5 | P5 | 10 | 1:59 | 4 | 0 | Deployed |
| 6 | Follow-up (counted): "The shipping lanes map is cut off at the bottom in the deployed app. Make every chart fit its card, then deploy it again." | 24 | 7:04 | 13 | 0 | min-h-0 on the shared card body; redeployed |
Run 2
| Turn | Prompt or answer | Words | Wall | Tool calls | Charts | What happened |
|---|---|---|---|---|---|---|
| 1 | P1 | 32 | 2:20 | 14 | 0 | Scaffolded dataapp and connected |
| 2 | P2 | 60 | 15:13 | 105 | 3 | The page; asked about provisioning to validate in the portal - "Skip browser validation for now" |
| 3 | P3 | 40 | 10:16 | 85 | 0 | Notes; ended waiting on the user |
| 4 | Answer: "Hold off for now. I'll ask you to deploy when I'm ready." | 12 | 0:09 | 0 | 0 | - |
| 5 | P4 | 23 | 4:55 | 27 | 0 | Scenarios |
| 6 | P5 | 10 | 6:01 | 42 | 0 | Deployed (it moved the template's Rayfin 1.33.1 pin itself) |
Run 3
| Turn | Prompt or answer | Words | Wall | Tool calls | Charts | What happened |
|---|---|---|---|---|---|---|
| 1 | P1 | 32 | 2:24 | 13 | 0 | Scaffolded dataapp and connected |
| 2 | P2 | 60 | 18:28 | 102 | 4 | The flow map failed its layout check and was swapped for a network diagram; the agent asked again and hosted the flow map. Asked about provisioning - "Skip browser validation" |
| 3 | P3 | 40 | 13:12 | 99 | 0 | Notes |
| 4 | P4 | 23 | 4:35 | 27 | 0 | Scenarios |
| 5 | P5 | 10 | 2:12 | 5 | 0 | Deployed |
The rubric (fixed before the series; a changed rubric is a new series). Hard checks: H1 the data app scaffold; H2
setup_fabric_app before the first chart; H3 the three intended charts are what the app hosts; H4 no interventions and
no turn stalled on a question; H5 every turn exits cleanly; H6 no wrangling in app code (row positions from the page,
mark geometry, a hand-built render payload, a slider rebuilt beside a what-if); H7a notes on marks through
annotations; H7b the scenario from the chart's own view state; H8 the cross-filter through the chart's click; H9
tsc -b, vitest and vite build pass. All three runs passed every check. (The fresh series in the loop chapter
adds H3b: the three charts first try, with no regeneration.)
The series, as spreads (prompts 1-4 and the stand-in's answers): prompts required 5 / 5 / 6; words typed 158 / 158 / 172; agent time 33 / 38 / 39 min; tool calls 231 / 241 / 247; output tokens 135k / 143k / 144k and cache reads 35M / 42M / 45M; generations 3 / 3 / 4; hand-written app code 1,345 / 1,348 / 1,350 lines.
Excluded runs, with the reason (a run hit by an environment outage is neither clean nor failed, and is replaced by a fresh identical run):
- Two early rehearsals on Sonnet 5.5: auto mode's safety check refused every shell command as "temporarily unavailable" (Appendix G).
- One where the harness killed the launching script before turn 3.
- One stopped mid-run when the test account was moved to the default model, so its inputs changed.
- In the deploy round, one run without our server is left out of the deploy count: our own harness had two apps of the same name in one workspace (Appendix H).
The standard route, three runs.
I asked a coding agent - Claude Code, on Claude Sonnet 5 - for the app in four prompts, each sent exactly as written, one per turn:
1. Create a new Fabric app called global-revenue from Rayfin's data app template, for the Fabric workspace "Global
Revenue Demo", and connect it to the semantic model "Global Revenue v2" in that workspace.
2. Build the main page for a regional sales director. I want a world map that compares revenue per capita with return
rate, a map of the shipping lanes between countries, and a revenue projection I can steer with a growth rate
against our target. Clicking a country should filter the other charts.
3. Let people click any country or shipping lane and leave a note on it. Save each note in the app's database with
who wrote it and when, and show it on that mark the next time anyone opens the page.
4. Let people save the growth rate they picked as a named scenario with a comment, and show everyone's saved
scenarios beside the projection.
A fifth, "Deploy it to the Global Revenue Demo workspace in Fabric.", came when I was ready. With Microsoft's Rayfin
plugin loaded, the agent read Rayfin's skill, scaffolded the app, connected it to the model and proved a live DAX query;
built the page with Microsoft's VegaVisual charts; saved notes in the app's own database with who wrote them and
when; and saved named scenarios. I ran the whole thing three times from an empty folder, with the same prompts and the
same setup each time.
What went well - and it's a lot. All three apps built, type-checked and passed their own tests. They themed well, handled loading and empty states, and followed the template's patterns. The agents wrote their own DAX against the model, and one caught a real trap in it: the model's what-if table only matches exact float steps (its 20% is stored as 0.19999999999999998). Clicking a country filtered the other charts. Each build took five or six prompts in all - the four, plus one or two one-click answers when the agent offered to provision or deploy something so it could check its work - and 38 to 43 minutes of agent time.
One surprise first: two of the three runs scaffolded Rayfin's blank app template, not the data app the prompt
names, because Microsoft's own getting-started skill says "Always use the bundled blankapp template". The agent then
reported that it had built from the data app template. It hadn't. Check the scaffold command yourself.
What didn't land was the presentation:
- The comparison map was bubbles. I asked for a world map that compares revenue per capita with return rate. Every run drew a bubble map - size for one measure, color for the other - and two of them on a sphere and graticule with no land drawn at all.
- The routes were straight lines between country centers - in one run, only the busiest 40 of the 291 lanes.
- The slider sat beside the chart. The steerable projection was a line chart with a slider built next to it, the model's what-if formula evaluated again in TypeScript.
- The maps were squashed. Deployed, one app's world map and lane map were drawn a few hundred pixels wide in full-width cards, and it took a follow-up prompt to make them fit.
Clicking around the deployed app turned up smaller things: tooltips that read just "true", legend clicks that selected the legend's text instead of filtering, a note that saved but was flaky to reopen from its lane, and no zoom or pan on either map. And the projection changed history: at a growth rate of +27%, every month back to 2019 sat above its actual. More on why in Part 2. None of the runs said it had drawn something other than what was asked.
The runs in full, and what happened when I asked each app to close the gaps: Appendix E.
With RegiaBI, three runs.
What was built, three times from an empty folder: a governed Fabric App over a Power BI semantic model, with a bivariate world map, a flow map of every shipping lane and a what-if projection you steer from inside the chart, click-to-filter across all three, notes left on marks and saved in the app's database, and named scenarios. It came from the same four prompts as Part 1 plus one sentence - "Use the RegiaBI charts server for the charts." - 155 words in all, on Claude Sonnet 5 in Claude Code, and it took five or six prompts and 33 to 39 minutes of agent time a run. The builder is still a command-line coding agent, with a setup list to go with it; that's where this stands today. Every prompt is below exactly as sent, Part 1 is the same prompts without our server, and Part 3 takes the same app onto the open web.
The only difference from Part 1 is the second sentence of prompt 2. Three runs a side, with identical inputs - the same prompts, setup, model, chart server build and stand-in - every number as min / median / max:
| Across 3 runs each | With RegiaBI | Part 1 (without) |
|---|---|---|
| Prompts required | 5 / 5 / 6 | 5 / 5 / 6 |
| Words typed | 158 / 158 / 172 | 150 / 151 / 161 |
| Agent time, prompts 1-4 | 33 / 38 / 39 min | 39 / 40 / 43 min |
| The three charts as asked | 3 of 3, every run | 0 of 3, every run |
The series converged: all three runs with our server passed every hard check of a rubric fixed before the series ran (the three charts asked for; no hand-built slider or render payload; notes on the marks; the cross-filter wired through the chart; build, type-check and tests green). One needed a second try at the flow map - the first failed its own layout check and was swapped for a network diagram, so the agent asked again. The extra prompts were one-click answers to the agent's own questions: in every run, an offer to provision the app in Fabric so it could check the page in the portal (answered "skip"), and once a "hold off, I'll ask you to deploy". The fifth prompt, "Deploy it to the Global Revenue Demo workspace in Fabric.", then deployed each app on its own. Per-run tables, every answer verbatim, the rubric and the excluded runs: earlier in this appendix.
The same three charts in every run:
- Bivariate World Choropleth - each of the 40 countries colored by revenue per capita and return rate through a 3x3 color key, so the places where the two disagree stand out.
- Origin-Destination Flow Map - all 291 shipping lanes as routes over a basemap, wider for more units shipped, with zoom and pan.
- What-if projection - monthly revenue against the target, projected forward from the last month at a growth rate and over a horizon you drag inside the chart.
Why those three: before choosing anything, each agent read the server's guide, ran each chart's DAX through
query_semantic_model so the whole result was profiled (40 countries, 291 lanes, 84 months), and asked what fits. In
these runs our server's instructions named bivariate maps, flow maps and what-if charts as examples of what the
catalog holds, and the agents asked for those types by name - I'd rather you knew that than assumed the server
volunteered them unprompted.
Clicking a country filters the lanes and the projection, and a second click clears it. A note left on a country or a lane appears as a badge on that mark - with the latest note as its title - the next time anyone opens the page. A saved scenario is the projection's own state, saved and applied again. At view time, what runs is the generated chart code and the open-source host: no key and no call to us, only DAX, as the signed-in reader.
In numbers: building the page took the runs without our server about 1.4x the time, 1.5 to 2.5 times the tool calls and about twice the agent's output (112-132k output tokens against 51-62k for that turn), and their finished apps carry about 1.5x the hand-written code (1,995-2,115 lines against 1,345-1,350) - for a bubble map, lines between centers and a slider beside a line chart. Per-chart files, lines and calls, and the code excerpts: Appendix D.
The wiring, as the first series measured it.
What the wiring costs with RegiaBI is short enough to print. This is the whole interaction code for the world map in the first run, trimmed of imports and the notes dialog:
const annotations = useMemo(() => buildCountryAnnotations(notes), [notes]);
const handleSelectRows = useCallback((rows) => {
const row = rows[0];
if (!row) { selectCountry(null); return; }
selectCountry({ code: row.CountryCode, name: row.Country });
}, [selectCountry]);
<BivariateWorldChoroplethChart rows={rows} {...size} options={options}
onSelectRows={handleSelectRows}
annotations={annotations}
onAnnotationClick={handleAnnotationClick} />
A click hands back the source rows; one key, the country code, goes into app state; and annotations puts a badge on
every mark that has notes, matched by the model's own key - a lane by both its ends - and kept there through every
redraw.
Without it, it was less work than I expected, and I should say so. Microsoft's VegaVisual emits click events too:
each run's cross-filter was a Vega-Lite selection plus about 20 lines that parse the event's predicates back into a
country, then shared state. Notes "on the mark" became a has-notes column joined into the data and drawn as a ring or a
pin: the mark says it has notes, not what they say. The notes turn took about as long either way, and saved scenarios
were quicker without us, because a slider you own is simple to save. So the wiring gap is real but modest. The widget
gap above is the big one. The maps article's "row identity" finding is check 5 here: a note finds its mark again only
if the mark carries the model's key. The evidence per check and the full measurements: Appendix D.
The clipped flow map.
And the one I'd most like you to take away: the clipped flow map. Deployed, the first run's flow map was cut off at the
bottom of its card. Microsoft's template skills do cover this - a "chart container height chain" recipe that warns
against exactly the pattern - and all six agents opened the app-design skill that points to it. Three opened the
recipe; the first run's agent didn't; and two that did still shipped squashed maps. Agents miss things, which is why I'd
rather rely on tools that avoid the gotcha first try - what RegiaBI is trying to become. The other half, told straight:
in the runs with our server, the faulty piece was our own chart panel, used exactly as we told the agent to use it. Its
fallback minimum height let an agent-written, fixed-height card slice the chart. One counted follow-up prompt fixed the
live app - "The shipping lanes map is cut off at the bottom in the deployed app. Make every chart fit its card, then
deploy it again." - in one seven-minute turn (the card's body needed min-h-0). Our panel now measures a box its own
chart can't inflate, gives each chart a minimum size and scrolls below it, so no card can do that to a chart again.
Every detour in full: Appendix H.
The fresh series that followed.
| Recorded series | Fresh series, attempt 3 | Fresh series, attempt 4 | |
|---|---|---|---|
| Chart server | 0.6.52 | 0.6.56 | 0.6.57 |
| Runs clean on every check | 3 of 3 (the rubric then) | 2 of 3 | 1 of 1 so far |
| Scripted prompts (with deploy) | 5 | 5 | 5 |
| Answers to the agent's questions | 1 / 1 / 2 | 1 / 1 / 1 | 3 so far |
| Follow-up prompts after deploy | 1 (the clipped flow map) | 1 layout tweak | 0 so far |
| Runs where a chart was regenerated or swapped | 1 of 3 | 1 of 3 | 0 of 1 so far |
Two things changed in how runs are scored, and both are worth knowing. The bar got stricter: the fresh series fails a run that regenerates a chart, where the recorded series didn't check. And two calls I made along the way: an agent's end-of-turn offer to deploy is a genuine choice, not a stall - Microsoft's template makes checking the app in the portal mandatory, and that needs a deployed app - so it's listed, not counted; and one visual tweak per run is allowed, because a reader who sees the page will ask for one (the run that used it had a projection scrolling beside its scenarios panel).
What changed between the series, all of it released: the flow map's basemap trim and the world maps' zoom wiring
became shared helpers the generated code calls instead of re-typing; the projection's target label fits at every panel
size; the flow map's key title fits; our chart panel can't be clipped by its card; check_chart checks every chart
locally before anything is deployed; the setup tool lines up every Rayfin package on one version and patches the
template's own failing test; the guide carries a recipe for notes on marks and the deploy traps; and the server
installs with no dependencies, so it's ready in seconds on a first run. The attempts in between failed honestly.
Attempt one lost two runs to the environment (a stopped launcher, and our test account running out of credits), and
its third run failed on the template's own test. Attempt two was stopped when a correct warning made the agent
hand-edit generated code. Attempt three had two clean runs and one that regenerated the flow map after its key label
escaped the chart.
What's still upstream - Rayfin's and Fabric's, documented and headed off rather than fixed: the data app template
pinning a Rayfin version too old to deploy; rayfin init's silent cancel and its follow-up demands; the data service
shipping off; Microsoft's own skill steering agents to the blank app; the tenant setting that needs Apply; a workspace
leaving Premium Per User blocking its models until a full refresh; two apps of one name sharing a Fabric item; and a
deleted app the project still remembers. Each is in Appendix A, with what heads it off.
The fourth attempt (chart server 0.6.57) finished three clean runs of three: five prompts each, three charts first try, every chart whole at the app's panel minimums, and first-try deploys.
K. The chart server, tool by tool
In a Fabric App the agent's path through the chart server is short:
setup_fabric_appadds the open-source chart host, a theme bridge, a measured chart panel and, when the app wants them, notes and scenarios backed by the app's own database. It never overwrites a file.query_semantic_modelruns a chart's DAX, renames theTable[Column]headers, saves every row and profiles them in one call, so the chart is chosen from the real data, not a sample.list_eligible_chartsreturns the chart types this data can honestly support, and says why one is refused.generate_chartwrites chart code that has already passed the server's render gates, plus a ready React component wired to the app's data.check_chartdraws each chart headless at its real panel size and says whether it draws, fits and clicks, before anything is deployed.
Sidebar: beside Vega-Lite, not instead of it. A chart in Microsoft's data app is a DAX query through Fabric's data SDK (
useSemanticModelQuery), its result bound into a Vega-Lite spec thatVegaVisualdraws. That's the one chart route the templates offer, and every skill they ship teaches it. Our charts use the same query hook - the same DAX, cache and sign-in - and replace only the last step: the table feeds a component that binds it to a chart generated for exactly those columns, drawn in D3 by@bicharts/chart-host. Vega-Lite draws bars, lines, scatters and the rest well, and Microsoft's own page says anything else "might require more iterations and validation". A two-measure color key, routes over a basemap and controls inside the chart are "anything else", and that's where ours start. Both can share a page, its queries and its theme.
L. Beyond Fabric: the same charts in other hosts
Because the charts are the same generated code in every host, the Fabric App is one place to put them, not the only one. Fabric Apps adds the governed data, the sign-in and the write-back; the charts don't depend on it.
| Route | What you get | What you lose | Who needs a license to view | What it costs to author |
|---|---|---|---|---|
| Fabric App (this article) | Governed data, Entra sign-in, notes and scenarios saved in the app's own database | It opens only for your tenant, inside the Fabric portal | A Fabric capacity (a trial, for us). Nothing from us | The MCP's authoring calls |
| Power BI report with LLM AI Charts and Maps for Power BI | The same three charts and click-to-filter, on any Power BI tier | Anything that writes back: saved scenarios and shared notes | Power BI's own licensing only: Desktop is free to build in, Pro shares it - no PPU, no Fabric capacity. A pinned copy needs no RegiaBI license to view | Generations in the visual |
| Web app | The same charts in React on a public page | Governed data and shared notes - notes stay in each visitor's browser | Nobody | The MCP's authoring calls |
In the broadest sense a Fabric App is a web app: a Vite + React single-page app with two Rayfin layers on top, data (DAX through the semantic-model connector) and backend (sign-in and the SQL database). Swap those two layers for the public CSVs the model reads and for storage in each visitor's browser, and the same page runs on the open web. The maps article's React demo and September's Blazor demo were the first two hosts off Power BI. For anyone, anywhere, that's the certain route.
Comments
No comments yet — be the first.
Leave a comment
Your email is never shown — we use it once to confirm the comment.