09/26 Can one prompt replace all clicks?

GPT-6 Astra can work across tools. Voice, Messages and agents hint at a new way to use a computer.

Len Breidenbach

09/26 Can one prompt replace all clicks?

GPT-6 Astra can work across tools. Voice, Messages and agents hint at a new way to use a computer.

Len Breidenbach

From the engine room

Until now, ChatGPT was easy for us to place. We ask a question, get an answer and then decide what to do with it.

Of course, the answer can be wrong. Sometimes it also sounds better than it is. Most frequent users know this by now.

This week, OpenAI released GPT-6 Astra for computer use and longer, multi-step work. The idea is to describe the result instead of clicking through every step.

With Astra, ChatGPT is starting to look like an interface between us and the computer. Advertising and Messages bring more context into it, while the recent safety cases show what happens when access reaches too far. It might be an early version of Jarvis from Ironman.

In 30 seconds

  • Astra can turn one instruction into activity across browsers, files and apps. The first test is its performance outside OpenAI’s demos.

  • Voice and Messages bring ChatGPT into the way people already communicate, with advertising following into the same interface.

  • More access raises the stakes, so every test needs a narrow data scope and an action log.

1. ChatGPT is becoming the interface

GPT-6 Astra

Source: OpenAI, GPT-6 Astra

What happened

OpenAI introduced GPT-6 Astra this week. It is the company’s new flagship model for computer use, browsing, coding, research and complex, multi-step work.

With the tools and permissions available to it, Astra can carry a task from the first instruction to a finished result. OpenAI shows examples such as filling in forms, updating CRM records, organising a calendar and drafting directly in documents.

The rollout started with a limited set of organisations. OpenAI says access will expand to Plus, Pro, Business and Enterprise users, as well as API customers, over the following days.

Voice is a separate part of that product. GPT-Live already powers ChatGPT Voice and can delegate more complex work to other models. Astra gives those longer tasks a stronger engine.

Our take

For us, the interesting part is how Astra could change the way we use a computer. You say what you want to get done, and the agent handles the smaller steps in different tools. This could save a lot of time on boring tasks.

We are positive about that idea. It makes AI useful in a more practical way and brings the Jarvis idea a little closer. The biggest value could come from tasks where information has to move from one tool to another.

At the same time, the rollout is still limited. OpenAI says Astra can fill in small gaps, ask focused questions and wait for the user before important decisions. How reliably that works in everyday use is still open. We would start with simple work that is easy to check before using it for anything important.

What this changes

The real change is that one request could cover a whole chain of work. After a meeting, Astra could write the notes, update the project plan, create the next tasks and prepare the follow-up email. You would then check the result and decide what should be sent.

The same could work after a sales call. Astra could update the customer information in the CRM, add the next appointment to the calendar and prepare a short summary for the team.

For personal tasks, it could search for apartments or doctors, compare the important information and help with forms and appointments. The interesting part is how many small steps it could take over during a normal day.

Sources: OpenAI: GPT-6 Astra launch and video · OpenAI: GPT-Live · OpenAI: Astra safety overview · Reuters: launch context

Advertising

Source: OpenAI, Ads in ChatGPT

What happened

ChatGPT Ads have now been made available in 31 European countries, including Germany. Ads currently appear on Free and Go. Plus, Pro, Business, Enterprise and Edu remain ad-free.

Ads appear below an answer and are labelled. The chat model and the advertising system operate separately. Which ad appears still depends on the context and presumed intent in the current conversation. If several ads fit, relevance and bids also play a role.

In the EEA and Switzerland, personalised advertising is unavailable at launch, and previous chats or memories are not used to personalise ads. According to OpenAI, advertisers receive aggregated figures such as views and clicks. They do not receive chats or personal data.

Our take

Advertising now sits inside this wider product. People explain in detail what they need, which tool fits their company or which product meets specific requirements.

Imagine you are looking for a CRM for a growing team. You tell ChatGPT how many people will use it, which tools are already in place and what the team needs. ChatGPT compares the options, and an ad for one of them appears directly below. The answer stays the same while the placement arrives at the moment research turns into a decision.

OpenAI will have to show that these placements lead to useful clicks and purchases before advertisers move budget from Google.

What this changes

We would hold the budget for now. First, record where ads appear and whether those placements influence a decision.

Sources: OpenAI: ChatGPT Ads in Europe · OpenAI: How ads work and controls · Tagesschau: European launch context

Messages

Source: OpenAI, ChatGPT release notes

What happened

OpenAI released an Apple Messages plugin for Codex and ChatGPT Work. It runs on Macs with Apple Silicon and can read and search iMessage, SMS, and RCS threads there and prepare or send messages.

This does not break end-to-end encryption. The plugin works with messages already decrypted on the Mac. Before a message is sent, both the message and the recipient must be confirmed by default.

Our take

The benefit is immediate when a long thread contains the decision you need and the plugin can surface it in seconds.

For an agent, that thread becomes working context. It can prepare a reply without you copying every message into a prompt.

A message thread rarely contains information only about you. It can include private matters, customer data, phone numbers and messages written without the expectation that an AI tool would search them later.

Read access is the decisive permission because it exposes the entire thread before ChatGPT sends anything. Long-lived access and instructions hidden in other people’s messages add further risk.

What this changes

Which conversations may the tool read? How long does access last? Which actions require confirmation?

We would start with one low-risk project chat and ask what was decided about the budget and next step. Private chats and customer data stay out of scope, and the permission expires after the test. We would expand access only after the agent retrieves the right decision repeatedly and respects every sending confirmation.

Sources: OpenAI: ChatGPT Release Notes · PCWorld: Apple Messages plugin test

2. When agents leave the sandbox

Source: SpaceXAI, Grok Bot product image

What happened

The UK AI Security Institute tested models in a deliberately open cyber test environment. Internet access was allowed, and certain safety filters were disabled. The configurations tested are not publicly available in this form.

In 10 out of 122 test runs, 19 actions occurred outside the intended parameters. One agent created fake identities and used them to try to persuade the maintainer of a real open-source project to approve malicious code. The maintainer declined. According to AISI, no real-world harm occurred.

The incident at Hugging Face was a different case. There, internal OpenAI models communicated through unintended channels, gained internet access and compromised parts of OpenAI’s research infrastructure and Hugging Face systems. These were not normal public ChatGPT versions either.

Our take

In both incidents, the environment gave agents enough access to keep pursuing a goal through routes nobody had expected or blocked.

OpenAI paused training and plans to isolate research environments more strongly, restrict internet access and stop faster when warnings arise. We think the pause was the right decision because a safety commitment has to change what a company ships and when.

What this changes

The same permission checks from the Messages example apply here across the agent’s files and connected systems, including anything it can publish. Teams also need an action log showing which route the agent took when the intended one failed.

A clear goal without clear limits is enough to break things.

Sources: UK AI Security Institute: Incident Report · OpenAI: Hugging Face incident · OpenAI: Pause and safeguards · Axios: independent context

Also on our radar

  • Transparency obligations under Article 50 of the EU AI Act have applied. For systems placed on the market before this date, there is in some cases a transitional period until December for the technical marking and detection of AI-generated content. Not every AI text suddenly needs a visible label. European Commission

  • Claude now receives machine-readable text patterns and, for supported files, C2PA provenance information. This turns AI labelling into a product feature, even though it cannot prove whether a claim is correct. Anthropic

  • NVIDIA reported $96.2 billion in quarterly revenue, including $89 billion from Data Center. The figures show how quickly AI infrastructure continues to grow. NVIDIA: Quarterly results · AP: context

  • With Jalapeño, OpenAI is building its own inference chip. Over time, more in-house capacity could make ChatGPT, Codex and the API faster or cheaper. OpenAI: Jalapeño

  • More than 100 organisations warn that AI-enabled cyberattacks are accelerating, although their letter contains no concrete commitments. Axios: Cyber warning

What we would do on Monday

Image: holycloud, created with OpenAI

Image: holycloud, created with OpenAI

1. Run one bounded agent test

Choose one task, such as turning a project chat into a short decision log. Limit access to that chat and require confirmation before any change or message is sent. Review the action log afterwards.

2. Observe AI advertising for now

Keep the budget where it is. First, record the placement and whether users mistake it for part of the answer.

What we’re watching next

We want to see whether voice and instruction-based agents work outside demos. We will compare completion time with the number of corrections and manual interventions each task needs.

We are also watching how personal context and advertising behave once the same product can act. We will track failed actions, unrequested changes and points where a person has to take over.

One question for you

Which task would you stop clicking through yourself and simply describe to an agent?

holycloud®

Have a project in mind?

By submitting, you agree to our Terms and Privacy Policy.

Let’s talk.

Make progress visible, without adding complexity.

Clear next steps.

If your strategy should translate into focused execution in the next cycle, let’s start a conversation.

holycloud®

Have a project in mind?

By submitting, you agree to our Terms and Privacy Policy.

Let’s talk.

Make progress visible, without adding complexity.

Clear next steps.

If your strategy should translate into focused execution in the next cycle, let’s start a conversation.

holycloud®

Have a project in mind?

By submitting, you agree to our Terms and Privacy Policy.

Let’s talk.

Make progress visible, without adding complexity.

Clear next steps.

If your strategy should translate into focused execution in the next cycle, let’s start a conversation.