When I wrote about GPT-6 Astra, the interesting question was: Can an agent take a messy problem, investigate it, and hand back something worth reviewing? The examples were deliberately substantial: bug investigations, BI recommendations, template-based reports, approval-gated record updates, and long-document synthesis.

Now there is a different question: Can you afford to do that every time the work shows up? Microsoft’s GPT-6.1 Sol announcement says GPT-6.1 Sol is generally available in Microsoft Foundry. It upgrades GPT-6 Sol, not Astra, and improves agentic coding, computer use, and professional work. Microsoft says it approaches Astra on its cited evaluations while offering a better capability-to-cost balance for agents that run all day. That is a positioning and evaluation claim, not a promise that Sol wins every task or that any particular deployment will be cheaper.

For a .NET team, I would not start with another one-off chatbot. I would start with the work that repeats: each PR, each support request, and each weekly review. No Python required; the patterns below use C# and Microsoft.Extensions.AI, and every sample prints the token usage you need to measure cost.

What changed since Astra?

A quick note on naming, because there are three models in play. GPT-6 Astra is the flagship I covered last time. GPT-6 Sol is the earlier Sol model. GPT-6.1 Sol is the upgrade to GPT-6 Sol. In the table, the Astra column summarizes how I used Astra; the GPT-6.1 Sol column describes what Microsoft says improved compared with GPT-6 Sol, not compared with Astra.

Question GPT-6 Astra GPT-6.1 Sol (changes vs. GPT-6 Sol)
Best starting point The most demanding reasoning and high-stakes reviews A candidate default for recurring production agents and complex workflows
Agentic coding Investigate bugs and propose changes for review Improved planning, editing, testing, and iteration across multi-tool workflows
Computer use Navigate approved interfaces with oversight Improved reliability across multi-step interface workflows
Professional work Synthesize, analyze, and produce review-ready work Stronger recurring analysis, drafting, and factual accuracy
Context and modalities Large text/image inputs Text and image inputs, text output, up to a 1M-token total context window; not a claim of more context than Astra

The comparison matters: 6.1 Sol is an upgrade to Sol; Astra remains Microsoft’s starting point for the hardest reasoning. Microsoft’s model reference lists a 1,050,000-token context window for both Astra and 6.1 Sol (maximum input 922,000; maximum output 128,000). A million tokens is capacity, not a suggestion to ship your whole repository on every request. Long prompts have real latency and cost consequences. Check the current catalog for your region and deployment options before building around any limit.

Foundry’s announcement lists Standard deployments across Global and US, EU, and APAC Data Zones; at launch it lists Provisioned Throughput in Global and US Data Zone. Serving location, capacity, and pricing are separate design choices. Check current availability and the Azure OpenAI pricing page for your own deployment rather than treating a launch-day price table as a quote.

Already running GPT-6 Sol?

In code, moving to 6.1 Sol is a configuration change: create a new deployment and point your deployment name at it. With Microsoft.Extensions.AI nothing else in your C# changes. Treat it as a model change anyway. Run your existing prompts and evaluation set against both deployments side by side, compare outcomes and token usage, and only then switch traffic. A model that reasons differently can spend a different number of tokens on the same prompt.

The .NET setup

Deploy gpt-6.1-sol from the Foundry model catalog and note two values from the deployment’s details page:

  • Endpoint: the resource endpoint, https://your-resource.services.ai.azure.com, not a Foundry project endpoint ending in /api/projects/….
  • Deployment name: the name you gave the deployment, which may differ from the model name.

Then give your identity access. The sample uses Microsoft Entra ID through DefaultAzureCredential, not API keys:

  1. Assign the Cognitive Services OpenAI User role on the resource to whoever runs the code: your own account for local development, a managed identity in production.
  2. Locally, sign in with az login (or Visual Studio / VS Code Azure sign-in). DefaultAzureCredential picks that up.
  3. Role assignments can take a few minutes to apply. A 401 or 403 right after assigning the role usually means “wait”, not “wrong code”.

Create the project. I compiled every sample in this post against the package versions below on .NET 10; pinning them keeps the code from drifting as Microsoft.Extensions.AI evolves. Any .NET 8+ SDK works.

The endpoint and deployment name live in .NET user secrets, which are stored in your user profile outside the project folder, so they never end up in source control.

A few choices in that setup are deliberate:

  • Timeouts and retries. Reasoning models can take a while on hard prompts, so the network timeout is raised and every call gets a cancellation token with an overall budget. The Azure client already retries transient failures (such as 429 and 503) with backoff, so you don’t need to wrap calls in your own retry loop. Do count those retries when you measure cost and latency.
  • UseFunctionInvocation() lets the model call the C# tools you register. It’s harmless for samples without tools, which simply never trigger it.
  • UseOpenTelemetry() emits a trace span per model call, including token counts, following the OpenTelemetry GenAI conventions. Nothing is exported until you add the OpenTelemetry SDK and an exporter (Azure Monitor, OTLP, or console); then the same data lands next to Foundry’s tracing.
  • Report prints input, cached-input, output, and reasoning tokens plus wall-clock time for each call. Those are the numbers you’ll need for the cost comparison later in the post.

The snippets below are alternatives to append after that setup, not three workloads to fire at once. Tool bodies and inputs are illustrative, not a live GitHub, browser, or CRM integration. Tool calling through UseFunctionInvocation() can invoke your C# methods: give those methods narrowly scoped permissions and validate inputs server-side. For a new deployment, also verify the supported API surface in the model documentation.

Use case 1: PR triage on every pull request

The Astra bug demo asked the model to correlate logs with a commit. Sol’s production pitch is to run a similar loop on every PR: inspect CI, flag likely breakages, propose a focused test, and let a human decide whether to merge. The interesting part is not generating a paragraph of review comments; it is knowing when evidence is insufficient.

The program reads CI output from a file when you pass one, and falls back to a sample otherwise:

Output – Case 1

Illustrative result: “Observed: the expired single-promo test failed; 203 others passed. Hypothesis: the resolver now accepts an expired code in one path. Next: reproduce that test locally and inspect expiry validation in the changed method. Do not merge until the failure is resolved.” No invented diff, no claim to have run tests.

Wiring it into CI

To run it on every pull request, put the program in its own project (say tools/PrTriage) and call it from your pipeline. Here is a minimal GitHub Actions workflow. It runs the tests, asks the model for triage only when they fail (that’s where it helps, and it keeps cost down), and writes the result to the job summary:

The triage step writes the endpoint and deployment name from repository variables into the runner’s user secrets, so the same Program.cs runs unchanged locally and in CI. azure/login with OIDC means no API key is stored in GitHub; the workflow’s federated identity needs the same Cognitive Services OpenAI User role as before. Notice what the model can’t do: it gets read-only test output, it writes only to the job summary, and merge rights stay with your branch protection rules. Treat PR content as untrusted input (a test name or log line can contain prompt-injection text), which is one more reason to keep the model’s output advisory. Passing the last 200 lines of output instead of the whole log keeps input tokens predictable.

Use case 2: Back-office changes with a real approval boundary

Astra illustrated “look up a record, then propose an update.” Sol makes the pattern worth testing for frequent billing-email changes or account reconciliation. This is function calling, not a demonstration of computer-use/browser automation. If you do add browser-based computer use later, isolate the browser, scope credentials, and require human confirmation before sensitive writes.

This time the model doesn’t return prose. GetResponseAsync<T> asks for structured output matching a C# record, so your code receives a typed proposal it can route into an approval queue:

The record must come after all top-level statements, so keep it at the bottom of Program.cs.

Illustrative result: ACC-4471 billingEmail: billing-old@northwind-retail.com -> finance@northwind-retail.com (pending approval). This is the real approval boundary: the model has only a read-only tool, and its output is data, a ProposedChange your service validates (is it a well-formed email? does the account exist?) and puts in front of an authenticated human. TryGetResult returning false is a normal outcome you should handle, not an exception. Read-only tools and authenticated human-operated write paths are a stronger boundary than asking a model to be careful.

Output – Case 2

Use case 3: Weekly document review without fake citations

Astra’s report and filing examples showed what long context can do. For recurring professional work, try Sol on a smaller, source-labeled set of documents: a vendor SLA and an internal incident note. Ask for an evidence table rather than a sweeping “everything looks fine.”

Output – Case 3

Illustrative result: “The SLA requires a claim within 30 days [SLA §4]. An availability incident is recorded on September 16 [Incident INC-204]. Unknown: whether a claim was submitted; the incident note does not say [Incident INC-204]. Follow up with the contract owner before the relevant deadline.” Before sending such a summary to anyone, verify each cited claim against the source and confirm the deadline calculation and time zone with the contract owner. Large-context access does not guarantee grounded answers.

For a weekly job, the cached count in the report is worth watching. If the system prompt and long reference documents stay identical between runs and come first in the prompt, repeated input can be served from the prompt cache, which is typically billed at a lower rate.

Choosing Astra, Sol, or Luna with evidence

Microsoft positions three models for different jobs: Astra for the hardest reasoning, GPT-6.1 Sol for production agents and complex recurring workflows, and Luna for high-volume preparatory and data tasks: think classification, extraction, and cleanup that feed a larger workflow. Treat that as a shortlist, not a routing rule baked into your product.

I would put a representative set of your own PRs, support requests, and document reviews through all three and record: correct outcomes, unsupported claims, tool calls, retries, human escalations, end-to-end latency, and total cost per successfully completed task. Include failed and repeated attempts in the cost.

Reasoning effort is the other dial. Microsoft.Extensions.AI exposes it as a provider-neutral option, so comparing settings is a loop:

Output – Reasoning effort

Check which effort levels your deployment accepts before relying on one; an unsupported value can be rejected by the service. Watch the reasoning count in the output: the reasoning documentation notes that reasoning tokens are billed as output tokens, so a short answer can still be expensive. If High doesn’t produce measurably better triage on your PRs, you’re paying for tokens nobody reads.

This post uses the Chat Completions API because that’s what GetChatClient targets. Azure OpenAI also offers the Responses API, which is built for multi-turn agent work with server-side conversation state. Microsoft.Extensions.AI can sit on top of either, so your IChatClient code stays the same, but options and behavior don’t always map one-to-one between them. Pick one for your evaluation and keep it fixed while comparing models.

Keep approval checkpoints, least-privilege tool identities, audit logs, and prompt-injection defenses in place regardless of which model wins. Foundry offers evaluation, tracing, monitoring, and safety controls, but they do not remove the need to test your tools and enforce authorization in C#.

Beyond a console app

The console app keeps the samples short. In ASP.NET Core, register the same pipeline once with dependency injection and inject IChatClient wherever you need it. Run dotnet user-secrets init and the same two set commands in the web project; ASP.NET Core loads user secrets automatically in the Development environment:

For anything a person watches in real time, such as a support agent waiting on a draft, stream the response instead of waiting for the whole thing:

Background jobs like PR triage and the weekly review don’t need streaming; the non-streaming call is simpler and gives you the usage numbers in one place.

Wrapping up

Astra was my “can this do the difficult work?” model. GPT-6.1 Sol is the model I would test first for work that happens constantly: PR triage, back-office review, and routine document analysis. Its announcement is about capability per task at production frequency, not about a bigger context window or magic autonomy. Keep the human approval button, log the token counts, measure your actual outcomes, and then choose the deployment that earns its place in your .NET stack.

Resources

Building AI features in C#? I write about practical, no-hype prompt engineering and Azure AI patterns for .NET developers. Check out Prompt Engineering for .NET Developers — free, no Python required. You can also subscribe for more posts.