Grok 4.7 Opens Longer Agent Runs: Self-Checking Still Needs External Acceptance
TL;DR
SpaceXAI releases Grok 4.7 through its API, Cursor and a gradual GitHub Copilot rollout. Longer reasoning and self-checks improve delivery only if the original requirements survive and acceptance stays independent.
Grok 4.7 is available in Cursor and through the API, with an emphasis on longer agent tasks. That upgrade contains a hypothesis requiring external evidence: spending more time checking its own work should reduce missed requirements at delivery. If engineers still have to restore the original requirements one by one, a longer run has not improved that part of quality. Official release Cursor announcement
The SpaceXAI announcement dates the event to 2026-09-21 but provides neither a publication time nor a timezone. This article is dated September 22 in Taipei, with a 12:53 reporting cutoff, and covers the previous day’s release within the 48-hour window. Neowin published at 13:40 EDT on September 21, or 01:40 on September 22 in Taipei. That establishes when coverage appeared, not when the model became available. Announcement date Independent reporting
The company describes a larger pretrained model and reinforcement learning strengthened for extended agent work. Training emphasizes tasks taking hours to complete; the vendor claims better self-checking and management of long context. These are its descriptions of training and capability, not acceptance results from ordinary projects. Training and capabilities
The API documentation lists a 500,000-token context window. Standard requests with prompts below 200,000 tokens cost $2 per million uncached input tokens and $6 per million output tokens. At the threshold, the entire request moves to the $4 and $12 rates. A long task accumulating substantial context cannot be budgeted solely at the lower rates. Context capacity also does not guarantee that every constraint remains in view. API specifications Pricing rules
Cursor has added the model and reports a CursorBench 4.0 score of 46.3% for Grok 4.7 at xhigh, compared with 40.4% for Grok 4.6 at high. Different reasoning settings mean those numbers alone cannot establish the improvement at equal computational effort. GitHub also announced a gradual Copilot rollout; an announcement does not mean every account can immediately select the model. Cursor evaluation Settings comparison Copilot availability
Keep completion criteria outside the agent
I would reserve this kind of extended capability for changes that require repeatedly tracing dependencies, such as an interface change affecting multiple modules. Fixing a clearly identified typo may not benefit in the same way from extra reasoning. The choice depends on how much searching and revision the task needs, rather than model rankings alone.
Suppose a team asks for a new sign-in flow while retaining support for older clients. An agent might finish the new flow but miss the compatibility requirement. If the product displays only the agent’s own declaration that it is finished, an engineer must rediscover which commitments were never checked. I would retain the original requirements and connect each to the actual changes and test results. Anything without supporting evidence should remain unconfirmed.
Self-checking can help an agent find errors, while acceptance criteria need to be fixed before execution. If the agent believes a test is outdated, it should explain the proposed change for a person to accept or reject. It should not relax the criteria itself and then declare success. This is a product-design judgment derived from the capability, not a claim that the hypothetical incident has occurred.
The release provides a usable model and integrations, but the available material does not establish acceptance rates for ordinary teams. Evidence that could test the opening hypothesis would show how many changes pass the agreed tests on the first attempt with the original requirements intact, and how many are returned because a requirement was missed. If omissions remain frequent after self-checking, longer agent runs need to be reconsidered.
The cover reuses this site’s Cursor and SpaceX brand image, not a performance chart for this release. Downloading an official launch image failed because network name resolution was unavailable.
Sources:
Related Articles
OpenAI Launches Dots: Ongoing Work Should Not Become a Growing Review Queue
OpenAI is rolling out dots to eligible paid users for work that continues between conversations. Read-only background research and limited memory controls make missed commitments and accumulated review work more useful tests than activity alone.
NVIDIA Expands Agent Safety: Runtime Permissions Do Not Guarantee Correct Work
NVIDIA launches its Open Agent Safety Platform, with OpenShell 0.1.0 enforcing file, network and tool permissions outside the model. That can contain mistakes, but more than 100 participating organizations do not establish deployment results or correct work.