Supported byAdvertisement
The new Copilot Studio home page makes it easy to start an agent. There is another label on the same page that deserves just as much attention: Uses Copilot Credits.
If your mental model is “we pay once the agent is published,” it is time to check which harness you are using.
With the GitHub Copilot harness, the work before publication matters too. That changes how I would plan a demo, a development environment and a production pilot.
I opened my ModernDevMgmt tenant to look at the current experience. The screenshots below are fresh portal captures from 28 September 2026, not Microsoft documentation images. I inspected existing demo agents without changing their configuration, running prompts or publishing them. Cost-control guidance is linked to Microsoft’s documentation; this was not a billing or enforcement benchmark.
Table of contents
The cost signal starts on the home page

The current home page in my demo environment. Agent and Workflow are marked GitHub Copilot, directly below Uses Copilot Credits.
Microsoft’s billing documentation says that creating, previewing, testing and evaluating experiences on this harness can consume credits. The charge reflects work involving models, tools and the harness. This scope is specific to the GitHub Copilot harness; it is not a statement that every historical Copilot Studio testing experience now has identical billing. Usage-based billing for this harness.
That is why I would ask for a development allowance before a workshop, not afterwards. It does not need to become a complicated procurement exercise. It does need an owner who understands what people are about to build and test.

“GitHub Copilot” here identifies the runtime
The name can create confusion. This is a Copilot Studio authoring and orchestration framework, not the GitHub Copilot service. Microsoft says that running these agents does not send customer data to that GitHub Copilot service. Standard agents remain available alongside the newer harness. Microsoft’s harness explanation.
For a pilot register, I would record the harness explicitly. “Copilot agent” is too broad. My fields would be:
- Agent name and purpose.
- Harness and environment.
- Business owner and technical owner.
- Knowledge sources and connected tools.
- Intended audience and publishing channel.
- Funding route and review date.
That small amount of context helps the next administrator answer a cost question without reopening every agent and reconstructing why it exists.
If you are still choosing a platform, my Foundry versus Copilot Studio guide covers the broader decision. This article starts after that choice: how do we build responsibly inside Studio?
Do not assume a trial environment means unlimited free testing
Microsoft’s current cost-management guidance states that developer and trial environments moved to usage-based billing on 1 September 2026 for this scenario. It also warns that historical non-billed consumption is only directional information, not a guaranteed future bill. GitHub Copilot harness cost management.
I would therefore avoid using an old screenshot of a free preview as the basis for a new workshop budget. Check the current entitlement and funding path for the environment you will use.
For the demo shown here, I am deliberately not claiming that the tenant has a paid allowance, a working stop rule or a particular monthly cost. Those need their own verification in the administrative cost views.
Review capacity, overage and the agent limit separately
The documented administration entry point is Power Platform admin center → Licensing → Copilot Studio. Review the environment’s allocated credits and whether additional tenant capacity or pay-as-you-go can fund continued use. Allocation alone is not enough to establish a spending boundary.
For the individual agent, Microsoft documents a monthly limit, notification threshold and Stop usage setting. Keep that separate from Azure budget alerts, which are visibility controls rather than automatic workload stops. Cost-control instructions.
I would capture those three decisions in one review record:
| Review item | Evidence I would keep | Owner’s decision |
|---|---|---|
| Environment allocation | The approved capacity and its scope | How much capacity is assigned to the pilot? |
| Overage route | Tenant-pool and billing-plan settings | Can the pilot continue using another funding source? |
| Agent limit | Limit, notification and stop behavior | When should this particular agent stop? |
If a control is unavailable, record that limitation. Do not silently replace “agent stops” with “someone receives an email.” Those are different promises to the team funding the work.
My live example is intentionally small
Here is the existing JR Demo – Rollout Coach – New UI Draft. It explains a fictional rollout policy. The Build page shows the model, a single knowledge file and the optional components around it.

The existing Rollout Coach draft, opened read-only. This screenshot is configuration evidence, not proof that the agent produced a correct answer or stayed within a cost limit.
I like this as a pilot shape because the acceptance criteria can be narrow. The agent should explain a rule, cite the policy version and refuse to claim that it changed a device. I do not need a collection of integrations to find out whether that basic experience is useful.
Before adding another tool, I would ask what new outcome it enables and how I will test that outcome. If the answer is only “the demo looks more impressive,” I would leave it out of the pilot.
Plan the tests before you start consuming credits
An unstructured test session can produce a lot of activity without teaching you much. I would prepare the questions and the expected evidence first.

The live Evaluate page explicitly flags Copilot Credit use. I opened the page but did not create or run an evaluation.
The following is a proposed test plan, not a report of executed tests:
| Test | What I would check | What would fail the test |
|---|---|---|
| A rule covered by the source | Correct policy version and relevant section | A plausible answer without source support |
| A rule absent from the source | Clear acknowledgement that the information is missing | Invented guidance |
| A request to deploy a change | Clear statement that this assistant does not execute changes | “Done” or “Applied” without execution evidence |
| A conflicting instruction in the question | The answer remains within the defined task | An unsupported exception or approval |
| A source update | The answer follows the reviewed replacement source | Continued reliance on outdated policy wording |
For each run, keep the prompt, outcome, error or retry, and reviewer decision. Record available consumption data after the relevant reporting delay. Do not turn a planned test matrix into a claimed pass rate.
Microsoft’s Preview experience supports interactive tests, including checking knowledge attribution and tool activity. The documentation also describes how configuration changes can affect subsequent turns in an active conversation. For a clean comparison, I would explicitly control the conversation state rather than comparing a fresh session with one full of earlier context. Testing an agent.
Compare accepted work, not the cheapest single response
A useful cost review asks whether the agent completed the task to the agreed standard.
Suppose two configurations answer the same policy question. One is concise but misses the exception. The other identifies the exception and cites the right section. Calling the first configuration cheaper does not settle the decision if someone must repair its answer.
I would compare total consumed credits for the test batch, accepted outcomes and human review effort. Keep failures in the denominator and retries in the consumption total. If you remove failed runs, the result becomes an advertisement rather than a useful engineering measurement.
This article contains no invented credit-per-answer numbers. Once the team has an approved allowance, the next step is to measure its own workload rather than borrow a price from an unrelated agent.

A stop rule needs a recovery plan
Microsoft documents that exhausted capacity can affect makers as well as end users, including authoring, preview and evaluation experiences. It also describes overage grace behavior. I would not market a small environment allocation as a precise, instant currency cutoff. Credit enforcement policy.
For an internal pilot, my recovery note would answer four questions:
- Who checks whether the interruption is capacity-related?
- Who may approve additional capacity?
- What should users do while the agent is unavailable?
- What evidence justifies restarting the pilot?
A one-page answer is enough. Without it, “Stop usage” can become an unexpected outage and an urgent request to remove the very control you introduced.
Before the next workshop
I would check the harness, confirm the environment and agree on the funding route before sending the first build prompt. Then I would prepare a small test set, record outcomes and review consumption together with quality.
That keeps the workshop focused on what we want to learn. It also makes the eventual budget conversation much less surprising.
Does unpublished mean no cost?
Not for the GitHub Copilot harness’s billable build and test activities. Treat development as part of the planned workload.
Should I disable every optional feature to save money?
I would start with the minimum useful scope, then add features when the test plan shows why they are needed. Fewer components also make failures easier to explain.
What should I bring to the person approving the pilot?
The use case, the target environment, the controls you intend to verify, the acceptance tests and the review date. That is more actionable than a request for “some Copilot Credits.”
Stay healthy, Cheers Jannik
New posts, straight to your inbox.
Hands-on guides on Intune, AI and Azure.
190+ guides · 5x Microsoft MVP · No spam, unsubscribe anytime · Privacy






