Skip to content
← Back to Insights

Grok 4.7 Opens Longer Agent Runs: Self-Checking Still Needs External Acceptance

AI Grok AI Agents Coding Product Design News

TL;DR

SpaceXAI releases Grok 4.7 through its API, Cursor and a gradual GitHub Copilot rollout. Longer reasoning and self-checks improve delivery only if the original requirements survive and acceptance stays independent.

Grok 4.7 Opens Longer Agent Runs: Self-Checking Still Needs External Acceptance

Grok 4.7 is available in Cursor and through the API, with an emphasis on longer agent tasks. That upgrade contains a hypothesis requiring external evidence: spending more time checking its own work should reduce missed requirements at delivery. If engineers still have to restore the original requirements one by one, a longer run has not improved that part of quality. Official release Cursor announcement

The SpaceXAI announcement dates the event to 2026-09-21 but provides neither a publication time nor a timezone. This article is dated September 22 in Taipei, with a 12:53 reporting cutoff, and covers the previous day’s release within the 48-hour window. Neowin published at 13:40 EDT on September 21, or 01:40 on September 22 in Taipei. That establishes when coverage appeared, not when the model became available. Announcement date Independent reporting

The company describes a larger pretrained model and reinforcement learning strengthened for extended agent work. Training emphasizes tasks taking hours to complete; the vendor claims better self-checking and management of long context. These are its descriptions of training and capability, not acceptance results from ordinary projects. Training and capabilities

The API documentation lists a 500,000-token context window. Standard requests with prompts below 200,000 tokens cost $2 per million uncached input tokens and $6 per million output tokens. At the threshold, the entire request moves to the $4 and $12 rates. A long task accumulating substantial context cannot be budgeted solely at the lower rates. Context capacity also does not guarantee that every constraint remains in view. API specifications Pricing rules

Cursor has added the model and reports a CursorBench 4.0 score of 46.3% for Grok 4.7 at xhigh, compared with 40.4% for Grok 4.6 at high. Different reasoning settings mean those numbers alone cannot establish the improvement at equal computational effort. GitHub also announced a gradual Copilot rollout; an announcement does not mean every account can immediately select the model. Cursor evaluation Settings comparison Copilot availability

Keep completion criteria outside the agent

I would reserve this kind of extended capability for changes that require repeatedly tracing dependencies, such as an interface change affecting multiple modules. Fixing a clearly identified typo may not benefit in the same way from extra reasoning. The choice depends on how much searching and revision the task needs, rather than model rankings alone.

Suppose a team asks for a new sign-in flow while retaining support for older clients. An agent might finish the new flow but miss the compatibility requirement. If the product displays only the agent’s own declaration that it is finished, an engineer must rediscover which commitments were never checked. I would retain the original requirements and connect each to the actual changes and test results. Anything without supporting evidence should remain unconfirmed.

Self-checking can help an agent find errors, while acceptance criteria need to be fixed before execution. If the agent believes a test is outdated, it should explain the proposed change for a person to accept or reject. It should not relax the criteria itself and then declare success. This is a product-design judgment derived from the capability, not a claim that the hypothetical incident has occurred.

The release provides a usable model and integrations, but the available material does not establish acceptance rates for ordinary teams. Evidence that could test the opening hypothesis would show how many changes pass the agreed tests on the first attempt with the original requirements intact, and how many are returned because a requirement was missed. If omissions remain frequent after self-checking, longer agent runs need to be reconsidered.

The cover reuses this site’s Cursor and SpaceX brand image, not a performance chart for this release. Downloading an official launch image failed because network name resolution was unavailable.

Sources:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.