Taswar Bhatti
The synonyms of software simplicity
ClaudeOpus5.5WithVSCode

In my Claude Opus 5.5 for C# developers post I called Opus 5.5 from my own code: the Anthropic.Foundry SDK, Entra ID, and the effort parameter. This time the code calling Opus 5.5 isn’t mine. It’s Claude Code, running in my terminal and in VS Code, pointed at my Foundry resource.

Microsoft already has a good step-by-step setup guide for Claude Code on Microsoft Foundry in VS Code, so I’m not going to repeat it. If you’ve never connected Claude Code to Microsoft Foundry, start there. This post is about what happens after it connects, when the model on the other end is Opus 5.5: how to make sure you’re actually talking to it, which dial to turn, and where the tokens go when you’re not looking.

Why Opus 5.5 Changes the Claude Code Setup

A quick recap of what shipped, and what each change means inside Claude Code:

  • It’s 40% cheaper than Opus 5. $4/M input, $20/M output, and $0.20/M cache reads. Claude Code sessions are mostly re-read context, so the cache price matters more than the headline price.
  • Medium effort is the default. Claude Code respects that: Opus 5.5 starts at medium, while most other models start at high. You can raise it per session.
  • Thinking is always on. There’s no thinking toggle to hunt for. Effort is the only dial, and MAX_THINKING_TOKENS does nothing on Opus 5.5.
  • Output is 30%+ faster. You’ll feel this in the VS Code panel on long diffs.
  • More refusals. The expanded safety classifiers apply in Claude Code too. If your repo is security tooling, expect some declines.

The catch: none of this matters if Claude Code isn’t using Opus 5.5. And by default, on Foundry, it isn’t.

The Short Setup (and the Line Most Guides Miss)

Install Claud Code Cli

Install Claud Code Cli

Deploy Claude Opus 5.5 from the Foundry Model Catalog to your Foundry resource, same as any other model. Give yourself the Azure AI User or Cognitive Services User role on that resource (either one is enough to call the model), and run az login. No API key: Claude Code falls back to the Azure credential chain when no key is set, so your az login session is the credential. If you plan to use the API Key then you will need to set your env key or the json file to have the ANTHROPIC_FOUNDRY_API_KEY rather.

Remember to install the Extension in VSCode

VSCode Claude

VSCode Claude

Now the part I care about. Instead of setx-ing environment variables or pasting them into every shell, put them in ~/.claude/settings.json. Both the CLI and the VS Code extension read that file, so you configure Foundry once:

The two model lines do different jobs, and you want both:

  • ANTHROPIC_MODEL makes Opus 5.5 the model for the session. This is the line most guides miss. On Foundry, Claude Code’s default model is Sonnet 4.5, not Opus. Setting only ANTHROPIC_DEFAULT_OPUS_MODEL remaps the opus alias, but you keep chatting with Sonnet until you switch.
  • ANTHROPIC_DEFAULT_OPUS_MODEL makes the opus alias (in /model, subagent definitions, and so on) resolve to your Opus 5.5 deployment instead of an older Opus you may not have deployed.

Use your actual deployment name if it isn’t claude-opus-5-5. ANTHROPIC_FOUNDRY_RESOURCE takes the resource name only, not a URL. If you need a private endpoint or custom domain, use ANTHROPIC_FOUNDRY_BASE_URL instead. Don’t set both.

In VS Code, install the Claude Code extension and add one line to your VS Code settings.json so it doesn’t push you toward an Anthropic login:

Then check it, in the terminal or the VS Code panel: /status

You’re looking for the API provider set to Microsoft Foundry, your resource name, and your Opus 5.5 deployment as the model. If the model line says Sonnet, the ANTHROPIC_MODEL line isn’t being picked up.

Dev takeaway: “Deployed in Foundry” and “used by Claude Code” are two different things. /status is the five-second check that tells you which one you’ve got.


Effort: The Dial You’ll Actually Touch

In the SDK post, effort was a property on the request. In Claude Code it’s a session setting, and there are three ways to set it:

  • /effort sets it for the current session: /effort high, /effort low, or /effort auto to go back to the model default.
  • The /model picker has an effort slider (left/right arrows). Whatever you pick there is remembered per model, so Opus 5.5 can keep its own setting.
  • CLAUDE_CODE_EFFORT_LEVEL in the environment or the env block wins over everything else, including /effort.

That last point is easy to trip over: put CLAUDE_CODE_EFFORT_LEVEL in your settings while testing, forget about it, and later /effort high appears to do nothing. For interactive work, leave the env var out and use /effort or the /model slider. Save the env var for scripted or CI runs where you want one fixed level. (The top-level effortLevel user setting also won’t apply to Opus 5.5, which is one more reason to use the per-model slider.)

Here’s how I map the levels to Claude Code work:

Effort What I use it for in Claude Code
low “Explain this file”, rename a symbol, write a commit message
medium (default) Everyday feature work, bug fixes, writing tests
high / xhigh Multi-file refactors, framework migrations, tricky concurrency bugs
max Rarely, and only for the session: /effort max when xhigh clearly isn’t getting there

Thinking tokens are billed as output, at $20/M. Raising effort for the whole day costs you; raising it for one hard problem and dropping back is cheap.


Where the Opus Bill Hides

Pinning everything to Opus 5.5 is easy. The surprise is how much else then runs on Opus 5.5 too.

Background tasks. Claude Code does small jobs behind the scenes, such as generating session titles. On the Anthropic API those go to Haiku. On Foundry they run on your primary model, which is now Opus 5.5, unless you deploy a Haiku model and point ANTHROPIC_DEFAULT_HAIKU_MODEL at it.

Subagents. When Claude Code fans work out to subagents (the Explore agent searching your repo, for example), each one uses the main conversation’s model unless something says otherwise. That’s Opus 5.5 for every file search. If you deploy a cheaper model, route subagents to it:

Again, those are deployment names, so use yours. Opus 5.5 does the planning and the edits in the main conversation; a cheaper model does the searching and reading.

Only reference models you actually deployed. Foundry has no startup model check, so Claude Code won’t warn you about a typo or a missing deployment when it launches. You find out mid-session. While drafting this post, a subagent in my own session died with:

It had tried a model my resource didn’t have. Everything else kept working, which is exactly why it’s easy to miss. If you only deployed Opus 5.5, leave the Sonnet and Haiku lines out and accept that everything runs on Opus.

Dev takeaway: with one deployment, every token is an Opus token. That can be fine, since Opus 5.5 is cheaper than Opus 5, but make it a decision rather than a surprise.


Make the $0.20 Cache Reads Count

Cache reads at $0.20/M are the best part of the Opus 5.5 price sheet, and Claude Code is a cache-heavy workload: every turn re-sends your instructions, CLAUDE.md, and the conversation so far. Caching is on automatically. Three things decide whether you actually get those cheap reads:

  • The default cache lifetime on Foundry is 5 minutes. Step away for coffee, come back, and the next turn re-writes the whole context at full price. For long sessions with gaps, set ENABLE_PROMPT_CACHING_1H to 1 in the env block. One-hour cache writes are billed at a higher rate than 5-minute writes, so this pays off for long, stop-and-start sessions, not quick ones. (Recent Claude Code versions also have CLAUDE_CODE_PROMPT_CACHE_TTL=1h, which applies only to the main conversation.)
  • Pick your model at the start and stay on it. Switching from Opus 5.5 to Sonnet and back mid-session throws away the cache each time.
  • Set up MCP servers before you start. Some Azure-hosted deployments reject Claude Code’s tool search, so it loads all MCP tools up front. Then connecting or removing an MCP server mid-session resets the cache.

Refusals Show Up in Your Editor Now

In the SDK post I made refusal handling a required pattern, because a refusal is a successful HTTP response with no answer in it. In Claude Code you don’t write that handler, but you’ll still see the result: Opus 5.5 declines more requests than Opus 5 in the biology, cybersecurity, and reasoning-extraction categories.

If you work on security tooling (scanners, fuzzers, detection rules), you’ll occasionally hit a decline on a request that looks reasonable to you. Rephrase the request with the defensive context, or do that piece by hand. Don’t try to engineer around the classifier; on a work resource, that’s a conversation for your security team, not a prompt trick.


Checking What You Actually Spent

Inside Claude Code, run /usage (/cost is an alias). On Foundry you get the session’s token counts, a prompt cache line, and an estimated dollar cost at list price. That estimate is great for “did turning effort up just double my session?”, and the cache line tells you whether the 5-minute or 1-hour setting is actually working.

It is not your bill. The real numbers live in Azure Cost Management for the Foundry resource, at whatever price your agreement gives you. Anthropic’s usage dashboards don’t see Foundry traffic at all. For per-developer numbers across a team, Claude Code can export usage through OpenTelemetry. Tagging the Foundry resource (team=…, env=dev) makes chargeback easier.


When Something’s Off

These are the Opus 5.5-specific problems I’d check first:

Symptom Likely cause Fix
/status shows Sonnet, not Opus 5.5 Only ANTHROPIC_DEFAULT_OPUS_MODEL is set Add ANTHROPIC_MODEL with your Opus 5.5 deployment name
model … is not available on your foundry deployment An alias or subagent points at a model you didn’t deploy Deploy it, or remove that line from env
/effort seems to do nothing CLAUDE_CODE_EFFORT_LEVEL is set and overrides it Remove the env var for interactive use
Bill higher than /usage suggests per session Background tasks and subagents running on Opus 5.5 Deploy Haiku/Sonnet and set the Haiku and subagent model lines
Every turn after a break is expensive 5-minute cache lifetime expired ENABLE_PROMPT_CACHING_1H=1 for long sessions
VS Code panel asks you to sign in to Anthropic Login prompt not disabled “claudeCode.disableLoginPrompt”: true
401 / 403 Missing role, or az login in the wrong tenant Azure AI User or Cognitive Services User on the resource; az login –tenant <id>

One more: there’s no /logout on Foundry. To switch accounts or tenants, change your az login.


Bottom Line

Opus 5.5 is the first Opus worth testing as your all-day default in Claude Code. It’s cheaper, faster, and medium effort is enough for most work. But “default” has to be deliberate on Foundry. Set ANTHROPIC_MODEL, confirm it with /status, decide whether background tasks and subagents should really run on Opus, and turn effort up only for the problem that needs it.

Configure it once in ~/.claude/settings.json, and the CLI and VS Code both pick it up. Then keep an eye on /usage for a week before you roll it out to the team.

Building AI features in C#? I write about practical, no-hype prompt engineering and Azure AI patterns for .NET developers. Check out Prompt Engineering for .NET Developers — free, no Python required. You can also subscribe for more posts.

ClaudeOpus5.5forCSharpDevelopers

Anthropic’s newest model, Claude Opus 5.5, just landed — meaning if you’re already building with Claude, you can call Anthropic’s most capable coding-and-reasoning model with a resource API key. Here’s what’s actually new, and how it looks from C#.

What Actually Shipped

A few facts worth knowing before you touch any code:

  • 40% cheaper than Opus 5. Input tokens at $4/M, output at $20/M (vs. $5/$25), with cache reads at just $0.20/M — a 60% drop that hits hard for agentic workloads where re-read context dominates the bill.
  • Medium effort is now the default. Anthropic reports Opus 5.5 matches Opus 5’s quality at noticeably fewer output tokens when running at the medium effort level. You control reasoning depth with the effort parameter (low → max) instead of the old thinking toggle.
  • Adaptive thinking is always on. You can no longer disable thinking. The effort parameter is your one dial for reasoning depth, latency, and cost.
  • 30%+ faster output. Opus 5.5 generates text noticeably faster than Opus 5 — meaningful when sessions run for hours unattended.
  • Sharper vision. Reads dense charts, diagrams, and screenshots more precisely, so a lot of prompt-side image preprocessing you may have built for earlier models becomes unnecessary.
  • New safety classifiers. Opus 5.5 will decline more requests than Opus 5 did (biology, cybersecurity, and reasoning-extraction categories), which means your error handling needs to check for refusals explicitly now.

Why This Matters for .NET Developers Specifically

The use cases Anthropic is calling out map directly onto real .NET workloads:

  • Long-running agentic coding. Overnight refactors, codebase-wide migrations, and multi-step debugging that would have taken engineering teams days — Opus 5.5 handles these with fewer turns and lower cost.
  • Enterprise knowledge work. Financial analysis, contract review, research synthesis across large document sets, and professional-grade spreadsheets and presentations.
  • Vision-powered workflows. Analyzing screenshots, technical diagrams, and PDFs directly — no separate OCR or vision model pipeline.
  • Cost-efficient scaling. The 40% cost reduction plus cheaper cache reads means you can run Opus 5.5 at larger scale without the bill spiraling — especially for agentic workloads where context per request has grown 2.6x in the last six months.

One thing worth knowing up front if you’re coming from OpenAI models in Foundry: Claude doesn’t use Azure.AI.OpenAI or Microsoft.Extensions.AI‘s OpenAI connector. It has its own official SDK — Anthropic, plus a Foundry-specific Anthropic.Foundry package for Azure-native authentication. As always: no Python, no notebooks — just C# and dotnet run.

Getting Started

Claude ships its own official .NET SDK, and Foundry gets its own package on top of it — this is not the Azure.AI.OpenAI pattern you might already have wired up for GPT models in Foundry.

Deploy Claude Opus 5.5 from the Foundry Model Catalog to your Foundry resource — same process as any other model. You need your resource name and a Foundry resource API key, not a separate Anthropic API key. Replace the placeholders locally; local/key authentication must be enabled on the Foundry resource for this sample to work.

API-key authentication is used in this sample. Using Anthropic.Foundry and AnthropicFoundryApiKeyCredentials with the resource key and resource name. Never commit API keys to source control or include them in screenshots or logs. User secrets keep keys out of the project files but are not encrypted and are for local development only. For production, use a managed secret store or environment variables supplied securely by your hosting environment.

Five Use Cases in One Program.cs

The explicit Anthropic reference pins the Messages API types separately from the Foundry authentication adapter.

Replace the generated Program.cs with the C# blocks below, in order: shared setup first, then Use Cases 1–5. They form one program, not five standalone top-level snippets. All using directives and executable startup code belong in this first block; subsequent blocks declare local functions. No other C# source files are needed. The console project and NuGet packages are still required.

The program runs case 1 by default. Pass 1–5 to select a case, or all to run them sequentially. For example, dotnet run — 2 runs the effort example, and dotnet run — all runs the full demo. Live calls incur charges. Case 3 is an offline refusal-handler test; case 4 skips if no screenshot is present; case 5 is a bounded, read-only agent demonstration, not an unattended repository editor.

Here’s the shared client setup and entry point:

Use your actual deployment name if it differs from claude-opus-5-5; set AZURE_AI_FOUNDRY_DEPLOYMENT through user secrets or an environment variable. Set AZURE_AI_FOUNDRY_RESOURCE and AZURE_AI_FOUNDRY_API_KEY for that same resource, with local/key authentication enabled. Compiling does not verify deployment availability, key validity, resource authentication settings, or model features.

Now append each use-case block below to that same file. Function-local parameters and message variables can safely reuse their names.


Use Case 1: A Basic Call — Same Messages API Shape

Whether you’re reviewing code, summarizing test failures, or asking a question about a codebase, the Messages API shape stays the same. Opus 5.5 just handles longer context and more complex reasoning better than its predecessor:

Dev takeaway: switching between direct Anthropic API and Foundry is basically a client constructor change — your request/response code stays identical, which matters if you want to keep a provider-agnostic AI layer.

Case 1 – Sample Output


Use Case 2: Effort — The New Cost/Quality Dial

This is the biggest operational change from Opus 5. Thinking can no longer be disabled. Instead, the effort parameter controls how much the model reasons through a problem before answering. The default is medium — Anthropic reports it matches Opus 5’s quality at noticeably fewer output tokens.

Use case: a build-pipeline bot that triages CI failures at low effort by default, and only escalates to high when the first summary is ambiguous — real cost control without a model swap.

Case 2 – Sample Output


Use Case 3: Refusal Handling — Now a Required Pattern

Handle the API’s refusal stop reason explicitly rather than assuming every successful HTTP response contains an answer. The shared PrintResponse function already does this for every live call. To exercise that branch reliably without trying to provoke a refusal, use this synthetic response fixture. It is a local handler test, not a model request or a complete API response.

Dev takeaway: a refusal can come back as a normal HTTP 200, not an exception. Always check StopReason; surface a clear message or escalate for review rather than automatically routing around the refusal. Authentication failures, throttling, and transport errors are separate failures handled at the entry point.

Case 3 – Sample Output


Use Case 4: Vision — Reading Screenshots Without a Preprocessing Pipeline

Opus 5.5’s sharper vision means you can send raw screenshots and get accurate analysis without a separate OCR step:

For example, dotnet run — 4 “C:\screenshots\error-dialog.png” selects this case. Relative paths are resolved from the current working directory. Only send screenshots you are authorized to share with the model; remove secrets and personal information first.

Use case: a support-ticket triage tool in your backend that reads user-submitted crash screenshots directly, skipping a separate OCR step.

Case 4 – Sample Output


Use Case 5: Long-Running Agent — The Overnight Refactor

Long-running coding work requires more than a single request: the application must preserve conversation history, execute allowed tools, and return each result with its matching tool-use ID. This runnable example demonstrates that loop with a read-only, in-memory sample repository and an eight-request limit. It asks for a migration plan, not file edits; no shell commands or model-generated code are executed.

That’s the final C# block: your single Program.cs now contains the startup code, the response helper, and all five use cases.

Validation scope: the original combined code was compiled with .NET SDK 10.0.302 and the package versions listed above. The offline refusal fixture, argument handling, missing-image behavior, and invalid-PNG guard were exercised locally. Live Foundry requests and the multi-turn agent exchange were not tested against an Azure deployment; those still require a valid Foundry resource API key, local/key authentication enabled, and a compatible model deployment. This code check does not verify the article’s pricing or model-performance claims.

Case 5 – Sample Output

Dev takeaway: the model proposes tool calls; your application decides what actually runs. A production overnight refactor additionally needs an isolated workspace, explicit approval for writes, a restricted test runner, checkpoints, cost limits, and telemetry. Direct Anthropic.Foundry calls do not automatically register this loop with Foundry Agent Service or provide end-to-end tracing; that integration requires separate setup.


Cost Optimization: Making Opus 5.5 Work Harder for Less

Opus 5.5 is already 40% cheaper than Opus 5, but there are practical steps to push costs down further on long-running workloads:

1. Pick the Right Effort Level

Not every task needs max effort. Use this guide:

  • low — CI triage, test summaries, quick b reviews
  • medium (default) — general-purpose coding, design tasks, knowledge work
  • high / xhigh — complex refactors, multi-file migrations, financial analysis
  • max — frontier-level benchmarks, extremely complex reasoning

2. Protect Your Cached Reads

Cache reads are just $0.20/M tokens — a fifth of what they cost on competing models — but they still add up at scale. As agentic coding has matured, organizations have shifted from asking developers to scale at all costs to asking developers to scale efficiently:

  • Pick your model at the start of a session rather than switching midway
  • Compact before you step away rather than after
  • If you’re on API keys or cloud providers, set the one-hour cache lifetime for long sessions
  • Forked subagents start from the parent’s cache instead of paying for the same context again

3. Measure Open-Ended Tasks

On well-scoped tasks, both Opus 5.5 and Opus 5 finish in about the same number of turns — the price cut is all you get. The gap is biggest on open-ended tasks, where a model can spend many turns on the wrong approach. Opus 5.5’s efficiency gains compound on harder, more ambiguous work.


Opus 5 vs. Opus 5.5 — Quick Comparison

Feature Opus 5 Opus 5.5
Input price / MTok $5 $4
Output price / MTok $25 $20
Cache read / MTok $0.50 $0.20
Default effort high medium
Thinking toggleable always on (effort parameter)
Output speed baseline 30%+ faster
Safety classifiers standard expanded (bio, cyber, reasoning)
Available on Foundry yes yes

Bottom Line

Opus 5.5 isn’t just a bigger model — it’s a shift from one-shot demos to sustained, hours-long work. Whether you’re refactoring a 680K-line bbase, analyzing financial filings, or building multi-step agents that run overnight, the 40% cost reduction plus the efficiency gains mean this is the first Opus model you should seriously evaluate for production workloads on Azure.

Deploy it from Foundry, configure your resource name and Foundry resource API key, and start with medium effort on your agentic workloads. You’ll see the cost savings immediately — and on harder, more open-ended tasks, the quality improvement compounds.

Ready to try it? Deploy claude-opus-5-5 from the Foundry Model Catalog and start with the code above. The Anthropic SDK for .NET is the only thing you need — no Python, no notebooks, just C#.

Source code: https://github.com/taswar/GptImage2.5-flare-sunburst-demo


Building AI features in C#? I write about practical, no-hype prompt engineering and Azure AI patterns for .NET developers. Check out Prompt Engineering for .NET Developers — free, no Python required. Also subscribe to my mailing list for the latest blogs, tips and tricks I share.

GPT-Image-2.5-Flare-Subburst

Picture a home design app. A customer types, “Show me a modern kitchen with dark cabinets and a large center island.” An image appears in seconds. They follow up: “Make it brighter, and add natural wood accents.” The app updates the image without redrawing the whole scene, without losing the layout that was already right. That’s the interaction GPT-image-2.5 — split into two models, Flare and Sunburst — is built for: fast, iterative image generation and editing that actually holds up across a real revision loop, not just a single lucky prompt.

Both models are now Generally Available from OpenAI or most providers. As always: no Python, no notebooks. Just Azure.AI.OpenAI‘s ImageClient, Microsoft.Extensions.Configuration, and dotnet run.

What Actually Shipped

  • Two models, one job split by speed vs. precision. Flare is the smaller, faster model for most production workloads — higher-quality images than GPT-image-2 at 50% lower latency, which is what keeps a user inside a live iteration loop instead of waiting on every change. Sunburst is for creative workflows that need greater precision and control — high-fidelity campaign assets and polished visuals where quality outranks speed.
  • More accurate editing. Update targeted elements while preserving the rest of the image — not a full re-roll every time you ask for a small change.
  • Stronger multi-turn editing. Refine through successive instructions without the image drifting or degrading as edits accumulate. This is the thing that actually breaks in a lot of image models: edit three quietly changes something edit one got right.
  • Better instruction following. More accurate interpretation of complex visual instructions, layouts, and stylistic direction.
  • Transparent background generation. Generate logos, product cutouts, UI assets, icons, and stickers directly with a transparent background — set background=”transparent” and output_format=”png” (PNG or WebP required; JPEG doesn’t support transparency).
  • Both GA, both production-ready. Unlike some of the other models in this catalog, there’s no preview caveat here — Flare and Sunburst are both fully supported for production workloads today.

Why This Matters for .NET Developers Specifically

The use cases map directly onto real .NET workloads:

  • Marketing and campaign production — generate and resize campaign assets from an approved creative direction, with legible headline text rendered directly in the image
  • Retail and e-commerce catalogs — produce and maintain large volumes of consistent product imagery at scale
  • Virtual try-on and personalization — let a shopper preview an item, then adjust color, style, or setting through successive edits without the result drifting
  • Education and training content — generate diagrams and illustrations that evolve as a lesson or explanation develops
  • Travel, real estate, and discovery apps — turn a description into personalized imagery a user can react to and refine

Both models sit behind the exact same ImageClient you may already be using for GPT-image-2 or GPT-image-1 — swapping in Flare or Sunburst is a deployment-name change, not a rewrite.

Getting Started

Deploy gpt-image-2.5-flare and gpt-image-2.5-sunburst from the Foundry Model Catalog to your project — same process as any other model. Grab your endpoint, and you’re ready to go.

Here’s the shared client setup every example below builds on:

Now let’s put both models to work on the scenarios that we are targeting.

Use Case 1: Home Design — Iterative Generation That Holds Its Shape

The scenario from the announcement, minus the voice layer: a customer describes a kitchen, sees it rendered, and asks for a change — and the change should read as a revision, not a brand-new image that happens to share a theme. Flare is the right choice here since this is a live iteration loop where responsiveness matters more than maximum fidelity.

Expected Output

kitchen-v1

kitchen-v1

The important detail is which API call each step uses: the first pass is a plain generation, but the revision is an edit against the actual first image — not a second generation call with an updated prompt. That distinction is what keeps the island and the camera angle consistent between v1 and v2 instead of getting a completely different kitchen that happens to also have wood accents.

Use Case 2: Travel Planning — Exploration with Flare, Final Asset with Sunburst

Exploring destinations and itineraries generates a lot of throwaway previews for every one image a traveler actually keeps. This is the two-tier pattern the Flare/Sunburst split is built for: cheap, fast previews during exploration, then a single higher-fidelity render once a direction is locked.

Expected Output

preview-0

preview-0

preview-1

preview-1

itinerary-day3-final-small

itinerary-day3-final-small

Paying full Sunburst price for every exploratory preview doesn’t make sense when most of them get discarded — this pattern only spends the higher-cost, higher-fidelity call once a customer has actually committed to a direction.

Use Case 3: Education — Diagrams That Evolve With the Lesson

A student discusses a concept with an AI tutor and gets diagrams that evolve as the lesson progresses — not one static diagram handed over up front. Flare regenerates a diagram each time the topic shifts, which only works if generation is fast enough to keep pace with a real conversation.

Expected Output

lesson-diagram-0

lesson-diagram-0

lesson-diagram-1

lesson-diagram-1

Wire this into a chat-driven tutoring UI and the diagram updates the moment a student asks a follow-up question — no separate “generate a diagram” button, no need for the student to describe what they want drawn in prompt-engineering terms.

Use Case 4: Retail Campaign Creative — Legible Text, Consistent Batches

Marketing and catalog work both come down to the same requirement: consistent output at volume, with headline text that’s actually legible — historically the weak point of diffusion models. Here’s a campaign banner from Sunburst (final quality matters) followed by a Flare batch of product catalog shots (volume matters more than maximum fidelity).

Expected Output

catalog-SKU-1001

catalog-SKU-1001

catalog-SKU-1002

catalog-SKU-1002

catalog-SKU-1003

catalog-SKU-1003

campaign-summer-sale-banner

campaign-summer-sale-banner

Open the campaign banner and check the headline text specifically — that’s the failure mode this model generation is explicitly built to fix. And because the catalog template is fixed with only the description swapped per SKU, lighting and framing stay consistent across the whole batch — swap the hardcoded array for a real product database and this loop becomes an overnight catalog-refresh job.

Use Case 5: Virtual Try-On — Multi-Turn Editing Without Drift

Virtual try-on lets a shopper preview an item and then adjust color, style, or setting — successive edits that need to stay consistent with each other, not drift with every change. This is Sunburst’s precision-editing strength: each edit builds on the last while the customer’s pose, framing, and identity stay intact.

Expected Output

customer-tryon-photo

customer-tryon-photo

tryon-navy

tryon-navy

tryon-navy-outdoors

tryon-navy-outdoors

Two edits deep, the customer’s face, pose, and the jacket’s shape are all still consistent — only the color and background changed, exactly as requested. That’s “stronger multi-turn editing” in practice: the kind of drift where edit three quietly changes something from edit one is exactly what this release is built to avoid, and it’s the difference between a try-on feature customers trust and one they give up on after two tries.

Where This Fits (and Where It Doesn’t)

Reach for Flare when:

  • You’re iterating live with a user in the loop and latency matters more than maximum fidelity
  • You’re generating at volume — catalogs, previews, personalization

Reach for Sunburst when:

  • The output is a final, ship-ready asset — a campaign hero image, a locked try-on result
  • Editing precision matters more than speed, especially across multiple successive edits

Don’t reach for either when:

  • You need pixel-perfect brand-guideline compliance with zero human review — treat generated output as a strong first draft, not an approval-free final asset
  • Your workload depends on non-text modalities or file attachments the image APIs don’t cover — check current capability docs before committing to a design

Wrapping Up

The pitch for GPT-image-2.5 isn’t “one bigger image model” — it’s a genuine two-tier portfolio: iterate fast with Flare, ship precise with Sunburst, and trust that a multi-turn edit chain won’t quietly drift by the third revision. For .NET developers, both models sit behind the exact same ImageClient you already know from GPT-image-2 — the new surface area is picking which model fits which step of your workflow, not learning a new SDK. Note: I had issues with the edit calls returing 404, thus I switched to HTTP Client

Source code: https://github.com/taswar/GptImage2.5-flare-sunburst-demo


Building AI features in C#? I write about practical, no-hype prompt engineering and Azure AI patterns for .NET developers. Check out Prompt Engineering for .NET Developers — free, no Python required. Also subscribe to my mailing list for the latest blogs, tips and tricks I share.

Azure 5 Highlights Thursday

This is the bonus content for the #124 edition of 5 Highlights Thursday — 7 highlights this week instead of 5. If you haven’t seen the main newsletter yet, check it out on LinkedIn.

1. What’s your favorite new GitHub Copilot feature?

A Copilot-focused video covering the workflow named in the title, useful for developers looking to apply AI assistance more consistently in day-to-day work.

2. What if an agent could refactor your agent?

A developer-focused look at AI agents, grounded in the specific build or platform scenario named in the title.

3. Event-driven agents, powered by MCP

A developer-focused look at Model Context Protocol patterns, grounded in the specific platform or workflow scenario named in the title.

4. A unified MCP layer with Toolboxes in Microsoft Foundry

A developer-focused look at Model Context Protocol patterns, grounded in the specific platform or workflow scenario named in the title.

5. When chatbots grow buttons: Building MCP apps with FastMCP

A developer-focused look at Model Context Protocol patterns, grounded in the specific platform or workflow scenario named in the title.

6. Building MCP servers with VS Code — Level up your MCP

A developer-focused look at Model Context Protocol patterns, grounded in the specific platform or workflow scenario named in the title.

7. What’s the deal with these IQs?

A current Microsoft developer update focused on the scenario named in the title, selected for this week’s digest based on recency and view count.


As always, please give me feedback on LinkedIn. Which bullet above is your favorite? What do you want more or less of? Other suggestions? Please let me know.

Download my Free Ebook Prompt Engineering for .NET Developers.

Last by not least, know someone who might be interested in this newsletter? Share it with them.

Subscribe to my Newsletter

Have a wonderful Thursday 😉

Taswar

MAI-Image-2.6

Every time a new image model ships, I ask the same unglamorous question: can I actually put this in a pipeline, or is it just good at one pretty hero shot? Most image models nail the demo and fall apart the moment you need the fiftieth consistent variant of the same product, or a single surgical edit that doesn’t quietly redraw everything else in the frame.

Microsoft AI just shipped two models that answer that question directly: MAI-Image-2.6 and MAI-Image-2.6-Flash, both now in public preview in Microsoft Foundry. They’re not “one model, two names” — they’re a genuine two-tier portfolio: one for maximum quality, one for maximum throughput, sharing the same capabilities underneath.

As always: no Python, no notebooks. Just C#, HttpClient, and a dotnet run.

What Actually Shipped

A few facts worth knowing before you touch any code:

  • MAI-Image-2.6 is independently verified as frontier-tier. No. 2 on the Arena text-to-image leaderboard at launch, ahead of Google’s Nano Banana family and Meta’s Muse Image. No. 3 on Arena’s image-editing leaderboard, and No. 1 for image editing on Artificial Analysis. This isn’t a marketing claim you have to take on faith — it’s ranked against the models you’re probably already comparing against.
  • Multi-reference editing. MAI-Image-2.6 accepts up to five reference images in a single request. Lock a product, a character, or a brand mark, and it stays consistent as you revise, resize, and reformat it. One approved concept travels across channel formats without drifting.
  • Web grounding. The model can pull real-world context from Bing Search instead of relying only on training data — meaningfully better accuracy for real places, objects, events, and subjects.
  • Explicit control over format and resolution. Adaptive aspect ratios, 1.5K output resolution, and levers over how much the model deliberates before generating — so the same model family serves both a fast iteration loop and a final production render.
  • MAI-Image-2.6-Flash is the throughput tier. More than 2x faster than GPT-Image-2-Medium and roughly 72–78% more efficient than GPT-Image-2, depending on the metric. Built for interactive apps, automated pipelines, and personalization at volume.
  • Pricing. MAI-Image-2.6 starts at $5/1M text input tokens, $8/1M image input tokens, $38/1M image output tokens. MAI-Image-2.6-Flash starts at $1.75/1M text input, $2.50/1M image input, $19/1M image output — genuinely cheap enough to generate per-user variants without flinching.
  • Enterprise-ready by default. Delivered as a Foundry Model sold directly by Azure — Entra ID auth, RBAC, key-based access, Azure data-handling commitments. Prompts and outputs are not used to train the models.

Choosing Between Them

This isn’t “pick the cheap one or the good one.” It’s a real division of labor:

  • MAI-Image-2.6 — final campaign assets, complex multi-reference edits, high-fidelity text rendering, anything where precision is the point
  • MAI-Image-2.6-Flash — rapid iteration, interactive experiences, personalization, high-volume generation, anything where throughput is the point

The intended workflow: explore and iterate with Flash, then render the final asset with MAI-Image-2.6 once the direction is locked. You’ll see exactly that pattern in the code below.

Prerequisites

Deploy both MAI-Image-2.6 and MAI-Image-2.6-Flash from the Foundry Model Catalog to your project — same process as any other model.

📝 Note: MAI-Image models aren’t compatible with the Azure.AI.OpenAI SDK’s ImageClient — they require the raw REST call against Foundry’s images endpoint, same as MAI-Image-2.5. The exact API surface is still settling in public preview, so confirm request/response shapes against the Foundry Model Catalog before shipping anything to production.

Here’s a small shared helper both examples below build on — auth token + a reusable POST helper:

Now let’s put both models to work on the actual use cases Microsoft is targeting.

Use Case 1: Marketing & Campaign Production — Legible Text, First Try

Text rendering inside generated images has historically been the weak point of diffusion models — garbled letters, wrong spacing, headlines that look right at a glance and fall apart on inspection. MAI-Image-2.6’s quality step-up specifically targets this.

Open the file and check the headline text specifically — that’s the thing to inspect closely, since it’s the failure mode this model is explicitly built to fix. If your marketing team currently routes every hero banner through a designer for text placement and kerning, this is the workflow to pilot first: generate ten headline variants overnight, let the team pick the one that needs the least manual cleanup.

Use Case 2: Retail & E-Commerce Catalogs — Consistent Imagery at Volume

Catalogs live and die on consistency: same lighting, same framing, same background across hundreds of SKUs. Here’s a batch pattern using Flash for the volume — this is exactly the “millionth personalized variant” scenario Flash is priced for.

Because the style template is fixed and only the product description changes, lighting and framing stay consistent across the whole batch. Swap this loop to read SKUs from your actual product database and you have an overnight catalog-refresh job instead of a photo studio booking.

Use Case 3: Brand & Creative Operations — Multi-Reference Editing

This is the capability that’s genuinely new in 2.6: up to five reference images in a single request, so a locked brand asset — a mascot, a product shape, a logo mark — stays consistent as you reformat it for different channels. Here’s a helper for a multi-reference edit call, followed by locking a product across three different channel formats.

Same mascot, same product, same brand colors — three different channel-native compositions, generated from two locked reference images instead of a designer manually recreating the layout three times. This is the actual “one approved concept travels across formats without drifting” pitch, and it’s the reason multi-reference editing is the headline feature of this release.

Use Case 4: Personalization at Scale — Flash in an Interactive Loop

Flash’s pricing and latency profile make per-user or per-segment generation viable inside a live request path, not just a batch job. Here’s a minimal pattern for a personalized greeting card generator — the kind of feature you’d wire into an app, not run overnight.

(Timings above are illustrative — measure against your own Foundry deployment region and load. The point of the pattern is that Flash’s latency profile makes this loop viable inside a request/response cycle, not that these exact numbers are guaranteed.)

Swap the hardcoded array for real user data from your app, and this same loop becomes a “generate a personalized asset on demand” endpoint behind an API controller.

Use Case 5: Image Cleanup & Surgical Editing

The other side of multi-reference editing is single-image precision editing — targeted object edits, replacements, inpainting, text updates, and artifact removal, without regenerating the whole scene. This is the same surgical-edit pattern from MAI-Image-2.5, still very much present in 2.6.

This is the unglamorous, high-volume use case that actually pays for itself: a warehouse team’s phone-camera product shots turned into catalog-ready images without a reshoot. The product itself is untouched — only the blur and background changed.

Where This Fits (and Where It Doesn’t)

Reach for MAI-Image-2.6 when:

  • You need final, ship-ready campaign assets with legible, precise text rendering
  • You’re locking a brand asset across multiple reference images and need it to survive reformatting
  • Editing precision matters more than generation speed

Reach for MAI-Image-2.6-Flash when:

  • You’re generating at volume — catalogs, personalization, high-throughput pipelines
  • Latency matters because the call sits inside an interactive user-facing path
  • You’re iterating on concepts before committing to a final MAI-Image-2.6 render

Don’t reach for either when:

  • You need pixel-perfect brand-guideline compliance without any human review — treat generated output as a strong first draft, not a final approval-free asset
  • Your workload depends on an SDK-level ImageClient integration today — you’re on raw REST calls for now, and the API surface is still settling in public preview

Wrapping Up

MAI-Image-2.6 and MAI-Image-2.6-Flash aren’t “the same model, one slower.” They’re a genuine quality/throughput portfolio: iterate fast with Flash, ship precise with 2.6, and use multi-reference editing to keep one approved creative direction consistent across every format your marketing or product team actually needs. If your team is currently paying someone to manually generate the fiftieth catalog variant or resize one approved hero image into six channel formats by hand — this is worth a real pilot, not just a demo.

Resources


Building AI features in C#? I write about practical, no-hype prompt engineering and Azure AI patterns for .NET developers. Check out Prompt Engineering for .NET Developers — free, no Python required. Also subscribe to my mailing list for the latest blogs, tips and tricks I share.

GPT-6 Astra for CSharp Developers

The next era of enterprise AI isn’t going to be defined by chat experiences. It’s going to be defined by how well a model can actually work for you — not just talk at you. GPT-6 Astra, OpenAI’s newest frontier model, is now generally available for all customers in Microsoft Foundry. Instead of optimizing for “answer this one prompt well,” Astra is built to take an open-ended challenge, reason through it in multiple steps, create a plan, and hand you a finished result.

That’s a meaningfully different design target, and it’s the kind of thing that matters once you move past demos and start building agents that have to survive contact with real workloads — where the interesting part isn’t the model call, it’s everything Foundry brings around it: identity, networking, governance, data handling, evaluation, compliance.

As always: no Python required, no notebook required. Just Microsoft.Extensions.AI and dotnet run.

What GPT-6 Astra Actually Is

A few things worth knowing before you touch any code:

  • Deliberate planning and decision support. Astra breaks a challenge into steps, evaluates options, communicates its recommendation, and identifies next actions for review — rather than just returning a single best-effort answer.
  • Polished, purposeful output. It applies context, templates, and quality standards throughout a workflow, aiming to produce documents, spreadsheets, presentations, and analyses that are ready for expert review, not rough drafts.
  • Execution across applications. With advanced tool use and computer use, Astra can interact with software on a person’s behalf, move between apps, and complete multi-step tasks with appropriate human oversight — including workflows that have no dedicated API.
  • Long-context understanding. Up to 1,050,000 total context tokens, so large repositories, filings, and multi-document knowledge bases can be reasoned about in one pass instead of chunked and reassembled.
  • Enterprise controls by default. Microsoft Entra identity and access management, encryption in transit and at rest, private networking, RBAC, content filtering, safety evaluations, and monitoring. Prompts and outputs are not used to train the models.
  • Generally available. Not preview — this is production-ready in Foundry today.

Why This Matters for .NET Developers Specifically

The enterprise scenarios Microsoft is calling out map directly onto real .NET workloads:

  • Software engineering — reproduce complex bugs, investigate likely causes, propose fixes, and prepare changes for developer testing and review
  • Business intelligence — build and refine dashboards in Power BI, compare data, identify trade-offs, and prepare insights to share
  • Professional work — produce documents, spreadsheets, and presentations that follow existing templates and business standards
  • Application workflows — update customer records, process forms, test websites, and work through approved interfaces where dedicated APIs are limited
  • Financial services — synthesize filings, market data, and internal research into an investment point of view, then draft client-ready materials in a firm’s house style
  • Customer experience — resolve tickets rather than route them: research the issue, act in CRM and billing, close the loop

And because Astra is a native OpenAI model in Foundry — not a partner/MaaS model — it slots in through the exact same AzureOpenAIClient + IChatClient pattern you already use for GPT-4o or GPT-chat-latest. No special client, no bearer-token workaround. Swapping Astra into your evaluation pipeline is a deployment-name change, not a rewrite.

Getting Started

Deploy gpt-6-astra from the Foundry Model Catalog to your project — same process as any other model. Grab your endpoint and deployment name, and you’re ready to go.

Use Case 1: Software Engineering — Deliberate Bug Investigation and Fix Proposal

This is the headline scenario: Astra reproducing a complex bug, investigating likely causes, and proposing a fix for developer review — not guessing from a stack trace alone. Let’s build a small agent that pulls recent error logs and the relevant code change history before it commits to a root cause, reasoning through both signals together instead of pattern-matching on the first plausible explanation.

Case 1 – Output

Notice Astra doesn’t stop at “here’s an exception” — it correlates the error with why it started happening (a specific recent change), proposes a concrete fix, states its confidence, and hands off a clear next action. That’s the “planning and decision support” Microsoft is describing, applied to something every .NET team actually deals with.

Use Case 2: Business Intelligence — Power BI Insight Synthesis

The second headline scenario is business intelligence: comparing data, identifying trade-offs, and preparing insights someone can act on. Here’s a pattern for feeding Astra a dataset summary — the kind of thing you’d pull from a Power BI dataset via the REST API — and getting back a structured, decision-ready recommendation instead of a paragraph you have to re-read three times.

Case 2 – Output

That last field is doing real work: Astra is explicit about where the data runs out and a human needs to step in, instead of confidently inventing a root cause it can’t actually support. Wire the structured fields straight into a Power BI custom visual, a Teams card, or an email digest — no regex-parsing a paragraph to extract “what do I actually do with this.”

Use Case 3: Professional Work — Template-Based Report Generation

The third scenario: producing documents that follow existing templates and business standards, polished enough for expert review rather than a rough draft. Here’s Astra generating a weekly status report against a fixed template structure — the kind of thing that normally eats twenty minutes of a project lead’s Friday afternoon.

Case 3 – Output

This is deliberately unglamorous, and that’s the point — Astra didn’t editorialize, didn’t invent a risk that wasn’t in the notes, and stuck to the exact template structure. That’s the difference between “ready for expert review” and “needs to be rewritten before anyone sees it.”

Use Case 4: Application Workflows — Acting Through Approved Interfaces

The fourth scenario is the one without a clean API: updating customer records, processing forms, and working through approved interfaces where a dedicated API is limited or doesn’t exist. Full computer-use automation is a Foundry-side capability with its own configuration, approvals, and monitoring — but the same tool-driven pattern applies at the code level. Here’s Astra deciding what action to take and why, with the actual system interaction going through a scoped, human-approved tool rather than the model touching anything directly.

Case 4 – Output

That “propose, don’t apply” boundary is the whole game here. The announcement is explicit about this: computer-use capability demands containment, with scoped credentials, approved resources, and human checkpoints for consequential actions. Your AIFunction tools are exactly where you enforce that boundary in code — a lookup tool that reads, and a propose tool that never writes without a human in the loop.

Use Case 5: Financial Services — Long-Context Synthesis Into an Investment Point of View

The fifth scenario leans on Astra’s up-to-1M-token context: synthesizing filings, market data, and internal research into a point of view, then drafting client-ready materials in a firm’s house style. You don’t need a chunking/retrieval pipeline for a single filing plus a research note — just pass the whole thing in.

Case 5 – Output

Every claim is tagged with its source — that per-claim citation discipline is exactly what you want before anything with “investment” in the name goes in front of a client, and it’s a direct product of feeding Astra the full source material instead of a lossy summary of it.

Where This Fits (and Where It Doesn’t)

Reach for GPT-6 Astra when:

  • You need deliberate, multi-step planning and decision support — not a one-shot answer
  • The output needs to be polished enough for expert review: documents, reports, structured recommendations
  • You’re automating work through approved interfaces or systems without a clean API, with human checkpoints on consequential actions
  • Your scenario genuinely needs long-context reasoning across large documents, filings, or repositories

Don’t reach for it when:

  • You need a lightweight, low-latency conversational bot — that’s a better fit for a chat-tuned model like GPT-chat-latest
  • The task is narrow, single-shot classification or extraction with no planning component
  • You’re not ready to build the governance layer (approvals, scoped credentials, monitoring) that responsible agentic/computer-use workflows require — Foundry gives you the tools, but you still have to configure and own that boundary

Wrapping Up

GPT-6 Astra’s pitch isn’t “smarter chat” — it’s “does more of the actual work and hands you something finished.” For .NET developers, that shows up as agents that investigate before they conclude, reports that follow your template without babysitting, and workflows that act through your systems with a human still holding the approval button. Pair it with Foundry Agent Service so that autonomy inherits identity, security, and lifecycle management rather than becoming its own liability — and it’s worth deploying gpt-6-astra next to whatever you’re running today and comparing the two side by side.

Resources

Grok 4.6 for CSharp Developers Agentic AI in Microsoft Foundry

Grok 4.6 — SpaceXAI’s latest frontier model — just landed in public preview in Microsoft Foundry as an Azure Direct Model. The headline isn’t “another big model dropped.” It’s that Grok 4.6 is built specifically for long-horizon, agentic work: planning across many steps, calling tools reliably, recovering when something goes wrong, and handing you a finished work product instead of a half-baked fragment you have to stitch together yourself.

That’s a meaningfully different design target than “answer this one prompt well.” And it’s exactly the kind of thing that matters once you move past demos and start building agents that actually have to survive contact with real workloads.

As always: no Python required, no notebook required. Just Microsoft.Extensions.AI and dotnet run.

What Grok 4.6 Actually Is

A few things worth knowing before you touch any code:

  • Frontier reasoning at value pricing. Grok 4.6 is positioned as the value-tier frontier option — frontier-class reasoning at a materially lower cost per task than comparable models. That matters the moment “reasoning agent” stops being a one-off demo and becomes something running continuously in production.
  • Selectable reasoning effort. You choose reasoning depth per call — low, medium, high, or xhigh (default high) — instead of paying maximum-reasoning cost on every single request regardless of whether the task needs it.
  • Long-horizon agentic execution. It’s designed to sustain complex, multi-step work — planning, tool calls, error recovery, and self-verification — with limited human babysitting.
  • Multimodal input. Text and images, so document-heavy, diagram-heavy, and screenshot-heavy workflows don’t need a bolted-on separate vision pipeline.
  • 200K token context window at launch. Solid for most agentic and document-analysis workloads — just set expectations up front if your scenario needs more.
  • Still preview. Validate against your own prompts, tools, and safety thresholds before anything production-sensitive touches it.

Why This Matters for .NET Developers Specifically

The use cases Microsoft is calling out map directly onto real .NET workloads:

  • Long-running agents — multi-tool orchestration with error recovery and planning, not a single tool call and done
  • Software engineering — multi-stage coding sessions, large refactors, and debugging across a real repository, not a single isolated function
  • Research and analysis — synthesizing dense source material into structured, decision-ready output
  • Enterprise knowledge work — drafting full documents, reports, and deliverables end to end, not just paragraph stubs

And it slots into Foundry the same way every other model does — same IChatClient abstraction, same deployment pattern. Swapping in Grok 4.6 next to GPT-4o or MAI-Thinking-1 in your evaluation pipeline is a config change, not a rewrite.

Getting Started

Deploy grok-4.6 from the Foundry Model Catalog (currently Global Standard deployment only) to your Foundry project, same as any other model. Grab your endpoint and deployment name.

Use Case 1: Long-Horizon Agent with Tool Calling and Selectable Reasoning Effort

Let’s build the thing Grok 4.6 is actually designed for: an agent that plans across multiple tool calls instead of answering from a single shot. Here’s a small deployment-readiness agent that checks a service’s health, checks its recent error rate, and only then decides whether it’s safe to proceed with a deploy — reasoning through the combination rather than just checking one signal in isolation.

Case 1 – Output

Notice the agent has to combine two tool results before it can reason to an answer — a healthy service with a spiking error rate is exactly the kind of situation where a single-signal check would give you a false “all clear.” That’s the long-horizon, multi-step reasoning Grok 4.6 is built for, not a gimmick.

Use Case 2: Software Engineering — Multi-Stage Code Review Across a Repository

The second headline use case is software engineering: multi-stage coding sessions and debugging across complex repositories, not a single isolated snippet. Here’s a pattern for feeding Grok 4.6 multiple related files and asking it to reason about the change as a whole, not file by file.

Case 2 Output

The interesting part isn’t reviewing one file — any model can do that. It’s reasoning across files: GetByIdAsync returns Order?, and the caller in OrderService uses it without a null check. That’s a cross-file bug a naive single-file review would miss entirely.

Use Case 3: Research and Analysis — Structured, Decision-Ready Output

Grok 4.6 is also positioned for turning dense source material into structured output you can actually hand to someone. Here’s a pattern using strongly-typed structured output instead of hoping the model formats things consistently.

Case 3 – Output

This is the difference between “the model wrote something that sounds like a recommendation” and “the model produced a typed object your application can actually act on” — route it to an approval workflow, log it, feed it into a dashboard, whatever your enterprise-knowledge-work pipeline actually needs downstream.

Use Case 4: Multimodal Input — Reviewing a Diagram or Screenshot

Since Grok 4.6 accepts text and images, document- and diagram-heavy workflows skip the separate vision pipeline entirely.

Case 4 – Output

Same IChatClient, same message pattern as every text-only example above — you’re just adding a DataContent alongside the TextContent.

Where This Fits (and Where It Doesn’t)

Reach for Grok 4.6 when:

  • You’re building agents that need to plan, call multiple tools, recover from errors, and run with limited human oversight
  • You’re doing large, multi-file refactors or debugging that require reasoning across a whole change set, not one file at a time
  • You need frontier-class reasoning at high volume, where cost per task actually matters
  • Your workload benefits from tunable reasoning effort instead of paying maximum cost on every call

Don’t reach for it when:

  • Your context needs meaningfully exceed 200K tokens — qualify that requirement up front
  • The task is a simple, low-latency single-shot response where low effort on a cheaper model is a better fit
  • You need a production SLA today — it’s still preview, so validate against your own prompts, tools, and safety thresholds first

Wrapping Up

Grok 4.6 in Foundry isn’t another “yet another model” announcement — it’s a genuine value-tier option for the agentic, long-horizon work that’s increasingly what “building with AI” actually means once you’re past the demo stage. Selectable reasoning effort means you’re not stuck overpaying for depth you don’t need on every call, and the fact that it drops into the same IChatClient abstraction as everything else in Foundry means adding it to your evaluation lineup costs you an afternoon, not a rewrite.

Source code at: https://github.com/taswar/Grok46MSFoundryDemo


Building AI features in C#? I write about practical, no-hype prompt engineering and Azure AI patterns for .NET developers. Check out Prompt Engineering for .NET Developers — free, no Python required. Also subscribe to my mailing list for the latest blogs, tips and tricks I share.

Azure 5 Highlights Thursday

This is the bonus content for the #121 edition of 5 Highlights Thursday — 7 highlights this week instead of 5. If you haven’t seen the main newsletter yet, check it out on LinkedIn.

1. Improving Performance in .NET Applications

A deep dive into performance improvement techniques for .NET applications — covering profiling, benchmarking, and optimization strategies that matter for production .NET workloads.

2. Stop Writing Prompts. Start Writing Specs.

A reframing of how to work with AI coding tools: instead of ad-hoc prompts, write structured specifications that give AI agents the context they need to produce consistently better results.

3. What’s New in Azure Firewall? | Latest Updates for Smarter Network Security

A walkthrough of the latest Azure Firewall updates — covering new features and improvements that make network security smarter and easier to manage for teams protecting Azure-hosted workloads.

4. Build Reusable Copilot Workflows with Skills and Agents

How to build reusable, composable workflows in GitHub Copilot using skills and agents — a practical look at making AI coding assistance scale across your team and projects.

5. Build Intelligent Agents with Work IQ, Foundry IQ, and Fabric IQ

A walkthrough of building intelligent agents using the IQ platform stack — Work IQ for enterprise data tasks, Foundry IQ for AI model integration, and Fabric IQ for data and analytics workflows.

6. Optimizing GitHub Copilot: Better Results, Fewer Tokens

Practical guidance on getting more out of GitHub Copilot by writing better context, reducing token waste, and structuring prompts so the model produces tighter, more accurate results.

7. Modern WinForms Development: How AI is Reshaping Your Approach

A look at how AI tools are changing WinForms development in 2026 — from AI-assisted code generation to modern patterns for working with legacy Windows Forms applications.


As always, please give me feedback on LinkedIn. Which bullet above is your favorite? What do you want more or less of? Other suggestions? Please let me know.

Download my Free Ebook Prompt Engineering for .NET Developers.

Last by not least, know someone who might be interested in this newsletter? Share it with them.

Subscribe

Have a wonderful Thursday 😉

Taswar

Azure 5 Highlights Thursday

This is the bonus content for the #120 edition of 5 Highlights Thursday — 7 highlights this week instead of 5. If you haven’t seen the main newsletter yet, check it out on LinkedIn.

1. Everything You Need to Know About the Latest in C#

A comprehensive walkthrough of what’s new in C# for 2026 — covering the latest language features and improvements relevant to .NET developers staying current with the platform.

2. Modernizing .NET Applications

Practical guidance on modernizing existing .NET applications — covering the tools, patterns, and strategies available for bringing legacy .NET code into the current platform era, including GitHub Copilot-assisted upgrade paths.

3. Explore the Future of ASP.NET Core & Blazor in .NET 11

A forward look at what’s coming in ASP.NET Core and Blazor in .NET 11 — covering new capabilities and the direction for server-side and client-side web development on the .NET platform.

4. SQL MCP Server: Bringing AI Agents to Your SQL Data

How the SQL MCP Server makes your SQL Server databases accessible to AI agents through the Model Context Protocol — letting agents query, reason over, and act on structured data without bespoke integrations.

5. .NET Developer Productivity with AI

How AI tooling — including GitHub Copilot and .NET AI Building Blocks — is being integrated into the everyday .NET developer workflow to accelerate coding, debugging, and app design tasks.

6. Building Intelligent .NET Applications: From AI Features to Production

End-to-end guidance on integrating AI capabilities into .NET applications — from choosing the right AI features to wiring them into production-ready .NET apps with proper tooling and patterns.

7. What’s new in vector indexing for Microsoft SQL | Data Exposed

An update on vector indexing improvements in Microsoft SQL for 2026 — relevant if you’re building semantic search or RAG applications that need to store and query embeddings at the database layer.


As always, please give me feedback on LinkedIn. Which bullet above is your favorite? What do you want more or less of? Other suggestions? Please let me know.

Download my Free Ebook Prompt Engineering for .NET Developers.

Last by not least, know someone who might be interested in this newsletter? Share it with them.

Subscribe to my newsletter

Have a wonderful Thursday 😉

Taswar

GPT-chat-latest (GPT-5.6 Sol) for C# Developers

Microsoft has a habit of quietly upgrading the model behind a stable endpoint name instead of forcing everyone to migrate to a new one. That’s exactly what just happened with GPT-chat-latest — it’s now built on GPT-5.6 Sol, and if you’re already pointing at GPT-chat-latest in Microsoft Foundry, you get the upgrade without touching a line of code.

If you’re not already using it, this is a good moment to start. Here’s what changed, why it matters, and — because I don’t do AI posts without something you can actually run — a working C# sample.

What Actually Changed

A few things worth knowing before you touch any code:

  • Same endpoint, better model underneath.. For gpt-5.6-sol You don’t select a new model name — you just get the improved behavior through the same integration path you already have.
  • More focused responses. Tighter formatting, more direct answers, and a clearer main recommendation instead of a wall of hedged possibilities.
  • Improved factual reliability. Fewer mistakes on the things that are easy to get subtly wrong — dates, numbers, sources, rules, and assumptions.
  • More consistent behavior. Whether the question is a one-liner or a genuinely deep multi-step task, the model handles both without feeling like two different personalities.
  • Multimodal by default. Text, image, and audio inputs with long-turn consistency — you’re not stuck bolting on a separate vision model for basic multimodal chat.

None of that is exotic. It’s the boring, unglamorous stuff that actually matters when you’re shipping a chatbot people rely on — not a demo you show once and never touch again.

Why This Matters for .NET Developers Specifically

If you’re building any of the following, this update is aimed at you:

  • Customer support and self-service bots — troubleshooting, product Q&A, multi-step processes grounded in your own knowledge base
  • Planning and knowledge-work assistants — breaking down objectives, reconciling constraints, producing briefs and recommendations
  • Multimodal conversational features — combining text and image context in a single chat flow
  • Retrieval-grounded assistants — multi-turn conversations that synthesize answers from documents you feed in, not just the model’s training data

And because it’s the exact same IChatClient interface you’re already using for GPT-4o, MAI-Thinking-1, or anything else in Foundry — there’s no new SDK, no new package, no new mental model. You point at the same deployment name and the improvements just show up.

Getting Started

Deploy gpt-5.6-sol from the Foundry Model Catalog to your project — same process as any other model. Grab your endpoint, and you’re ready to go.

A Multi-Turn, Retrieval-Grounded Support Assistant

Reasoning quality is nice, but the real test for a chat model is whether it holds context sensibly across a back-and-forth conversation, and whether it sticks to the facts you actually gave it instead of making something up. Let’s build a small support assistant that’s grounded in a product knowledge snippet, and push it through a multi-turn conversation.

Output

What I’m actually testing here: turn 2 depends on the model remembering what “that” refers to from turn 1 (clearing the cache), and turn 3 deliberately asks something the knowledge context doesn’t cover — a model with improved factual reliability should say “I don’t have that information” instead of confidently inventing a mobile-app answer. That’s the difference between a chatbot people trust and one that quietly erodes trust one hallucinated answer at a time.

Multimodal Input: Text + Image in the Same Conversation

Gpt-5.6-sol handles multimodal input natively, which matters for support scenarios where a user just wants to send you a screenshot instead of describing an error message character by character.

Output

Same IChatClient, same message list pattern — you’re just adding a DataContent alongside your TextContent in the same ChatMessage. No separate vision API, no separate client to wire up.

Where This Fits (and Where It Doesn’t)

Reach for GPT-chat-latest when:

  • You’re building conversational, multi-turn experiences — support bots, internal assistants, sales enablement tools
  • You need retrieval-grounded answers that stick to the facts you provide, not the model’s general knowledge
  • You want multimodal input (text + image) without adding a separate model to your stack
  • You want model improvements over time without maintaining multiple endpoint names in your config

Don’t reach for it when:

  • You need deep, extended multi-step reasoning as the primary workload — that’s a better fit for a dedicated reasoning model
  • You’re doing narrow, single-shot classification or extraction with no conversational component
  • Audio input is core to your scenario and you need to validate current format/latency support against your specific requirements before committing

Wrapping Up

The best kind of model update is the one where you don’t have to do anything. GPT-chat-latest running on GPT-5.6 Sol is exactly that: same endpoint, same IChatClient code, better answers underneath. If you’re building support bots, planning assistants, or anything that leans on multi-turn conversation grounded in your own data, it’s worth pointing your existing integration at it and seeing the difference for yourself.

Source code at: https://github.com/taswar/GptChat-Sol-Demo


Building AI features in C#? I write about practical, no-hype prompt engineering and Azure AI patterns for .NET developers. Check out Prompt Engineering for .NET Developers — free, no Python required. Also subscribe to my mailing list for the latest blogs, tips and tricks I share.

UA-4524639-2