Google released the Gemini 3 series of models, with the Gemini 3 Pro performing well in a number of benchmarks, with a score of 37.51 TP3T in the HLE test and 31.11 TP3T in the ARC-AGI-2 test.The Gemini 3 Deep Think performed even better, with a score of 411 TP3T in the HLE test and 45.11 TP3T in the ARC-AGI-2 test.Meanwhile, Grok 4.1 was released for free, with Claude models extended to Azure and Microsoft 365 platforms. 45.11 TP3T.Meanwhile, Grok 4.1 was released for free and the Claude model was extended to Azure and Microsoft 365 platforms.
- 此摘要由AI分析文章内容生成,仅供参考。

Recently, the Gemini 3 Pro's performance benchmark data was accidentally leaked, sparking widespread concern in the industry.

According to the leaked data, the Gemini 3 Pro performs well in several benchmarks:Score of 37.5% in HLE test, 31.1% in ARC-AGI-2 test, 2439 Elons in LiveCodeBench Pro test, 85.4% in Tau-Bench test, SimpleQA Verified test 72.1%The Gemini 3 Pro is the first of its kind in the industry. This excellent set of results demonstrates that the Gemini 3 Pro has reached an industry-leading level in a number of ways.
Despite performing slightly less well in the SWE-Bench Verified test, the Gemini 3 Pro still scored top marks in most of the key tests overall.This marks its breakthrough in artificial intelligenceThe
These outstanding performances of Gemini 3 Pro bode well for its great potential in future AI applications. Whether it is natural language processing, code generation, or complex problem solving, Gemini 3 Pro has demonstrated exceptional capabilities that promise to revolutionize several fields.
---

Google has released its latest AI model, Gemini 3, which is their smartest model to date, combining state-of-the-art reasoning capabilities, multimodal understanding, and powerful autonomous actions.
The Gemini 3 Pro topped the LMArena charts with 1501 Elo, and without tools it scored 37.5% on The Last Exam of Man, 91.9% on the GPQA Diamond test, and 23.41 on the MathArena Apex test. TP3T, and a factual accuracy score of 72.11 TP3T on the SimpleQA Verified test.These scores demonstrate the Gemini 3's excellence in a number of areas.
Availability and Pricing
- Gemini 3 is now available in the Gemini app (select "Think" mode), Google Search AI Mode (for U.S. Google AI Pro and Ultra subscribers, soon to be extended to all U.S. users), Vertex AI, Google AI Studio (free tier and rate-limited), Gemini CLI (for Ultra subscribers and paid API key holders, others need to waitlist), and Android Studio Otter.
- Pricing is $2 per million input tokens and $12 per million output tokens for cue words less than 200,000 tokens.Free access in Google AI Studio is subject to rate limits.
New Features and Performance
- Gemini 3's Deep Think mode scored 41.01 TP3T on The Last Human Exam, 93.81 TP3T on the GPQA Diamond test, and 45.11 TP3T on the ARC-AGI-2 test. which will be rolled out to Google AI Ultra subscribers in the coming weeks. After a security assessment.
- New features include a dynamic view and visual layout experiment generation interface, Gemini Agent for multi-step tasks, limited to Ultra subscribers in the U.S., and a redesigned application with My Stuff folder for easy finding of created content.
- Gemini 3 Pro scored 1487 Elo on WebDev Arena, 54.21 TP3T on Terminal-Bench 2.0, and 76.21 TP3T on SWE-bench Verified coding tasks.Integration is available at Cursor, GitHub, JetBrains, Manus, Replit, Cline, and Google Antigravity, which offers a free public preview on MacOS, Windows, and Linux.
Enterprise customers can access these features through Gemini Enterprise and Vertex AI, and eligible U.S. college students can get a one-year free trial of Google AI Pro, while Google AI Pro and Ultra subscribers will enjoy higher usage limits.
---

Recently, the Gemini 3 Pro's benchmark results were made public, showing that it performed well above expectations in a number of critical tasks. According to the newly released data, theHumanity's Last ExamThe Gemini 3 Pro scored 37.51 TP3T in the benchmark test, while theARC-AGI-2The test then achieved 31.1%.These results show that the Gemini 3 Pro is at the forefront of current artificial intelligence.

At the same time, the detailed documentation of the model has been accidentally leaked, which further aroused the attention of the industry.The outstanding performance of Gemini 3 Pro is not only reflected in the theoretical tests, but also in the wide range of application scenarios, including natural language processing, image recognition, and complex decision support, etc. The model can be used in a wide range of applications, including natural language processing, image recognition, and complex decision support.

- Humanity's Last Exam: 37.5%
- arc-agi-2: 31.1%
This series of breakthroughs will bring new impetus to the development of AI technology.
---

Recently, Grok version 4.1 was released and is free for all users.

The new version of Grok is significantly improved in several ways:
- It has been awarded the first place in the LMArena ranking, with an Elo score of 1483.
- Enhanced emotional intelligence to better understand and handle emotionally relevant tasks.
- Enhanced creative writing capabilities to provide users with richer and more varied creative options.
- Reduced illusionary phenomena and increased accuracy and reliability of generated content.
In addition, Grok 4.1 supports multiple platforms, including Web, X (formerly Twitter), iOS and Android, so users can access and use it anytime, anywhere.

Key Features:
- High emotional intelligence for sentiment analysis and customer service.
- Powerful creative writing aid for writers and content creators.
- Higher accuracy and reliability with fewer false outputs.
- Cross-platform support for easy use on different devices.
This update not only enhances the user experience, but also further solidifies Grok's leadership in artificial intelligence.
---

The Google AI developer team has unveiled Gemini 3 Pro, the latest generation of intelligent models that reach industry-leading levels of reasoning power and multimodal understanding.Gemini 3 Pro not only features powerful agents (agentic capabilities), and also has a unique ambient coding capability (vibe coding), to be able to better understand and process complex information.
These advanced features enable Gemini 3 Pro to excel in several application scenarios, such as natural language processing, image recognition, and cross-modal tasks. For developers, Gemini 3 Pro provides a rich set of API interfaces and development tools that facilitate rapid integration into existing systems.
In addition, Gemini 3 Pro supports a wide range of programming languages and frameworks, providing great flexibility for developers from different backgrounds. By using Gemini 3 Pro, developers can build smarter and more efficient solutions that push the boundaries of AI technology.
---

The latest news shows that Gemini 3 is now the best Vibe coding and proxy coding model. The power of this tool lies in its ability to build almost any type of project, whether it's a complex application or a highly interactive virtual environment.
Of particular note, Gemini 3 is particularly good at creating playable sci-fi worlds. Through the use of advanced shader technology, developers can achieve realistic visual effects that create an immersive experience. As a concrete example, users can explore a sci-fi world built with Gemini 3 at the following link:
https://t.co/T55LofFGN3
Key features include:
- Powerful Vibe Encoding Support
- Efficient Agentic Coding Model
- Rich Shader Library
- Suitable for a wide range of development scenarios
These features make Gemini 3 the tool of choice for developers, whether it's for game development, virtual reality, or other projects that require high-performance graphics processing.
---

Amazingly, Google's latest AI model, Gemini 3 Deep Think, outperformed its predecessor, Gemini 3 Pro, in several benchmark tests.
Specifically, on Humanity's Last Exam, Deep Think outperformed the Pro version by 411 TP3T; on the ARC-AGI-2 test, this difference was 45.11 TP3T.These data indicate that Deep Think has made significant improvements in its ability to understand and solve complex problems has been significantly improved.
Google re-establishes its leadership in AI with Deep Think. This could mark the dawn of a new era, especially in natural language processing and machine learning. Whether OpenAI will be able to catch up in the face of this challenge has become the focus of industry attention.
This breakthrough not only demonstrates Google's strong strength in technology development, but also provides new possibilities for future AI applications. With the continuous progress of technology, we are expected to see more innovative application scenarios, from intelligent assistants to complex decision support systems.
---

Update: Gemini 3 Deep Think scored 411 TP3T on the HLE (Humanity's Last Exam) test and 45.11 TP3T on the ARC_AGI-2 test.
These results show that Gemini 3 Deep Think outperforms its predecessor, Gemini 3 Pro, in a number of benchmarks.The HLE test focuses on evaluating the model's ability to deal with complex problems and reasoning tasks, while the ARC_AGI-2 focuses on the model's general-purpose AI capabilities.
Gemini 3 Deep ThinkNot only did it excel in these tests, it also made significant progress in the GPQA Diamond test. These achievements demonstrate its strong potential in the areas of natural language processing, reasoning and general artificial intelligence.
This breakthrough is important for advancing AI technology, especially in application scenarios that require advanced reasoning and understanding, such as intelligent assistants, automated question and answer systems, and complex data analysis.
---

The Gemini 3 Era has officially opened, a milestone that marks a new chapter in smart technology.

Gemini 3 Pro, one of the world's most intelligent models, will be widely used across multiple platforms and applications, including Google and its third-party products and services. The model has performed well in several benchmarks, particularly demonstrating its power in products such as AI Studio, Gemini API and Gemini App.
Key Features:
- Superior Performance: The Gemini 3 Pro performs well in a number of benchmarks.
- Wide range of application scenarios: supports various products and services such as AI Studio, Gemini API and Gemini App.
- Seamless Integration: Can be easily integrated into existing systems to improve efficiency and user experience.
The release of Gemini 3 Pro will provide developers and users with unprecedented innovation capabilities to help them break through in every area.
---
Anthropic has announced the full extension of the Claude model to multiple platforms, further driving its popularity in enterprise applications.
Azure customers now have access to Claude Sonnet version 4.5, Haiku 4.5 and Opus 4.1. Developers can use these models in Foundry with Claude Code for more efficient application development.
here's the thing., the Claude model is also integrated into Microsoft 365 Copilot and Excel's Agent Mode to provide users with a smarter office experience.
Anthropic has partnered with NVIDIA and Microsoft, making Claude the only cutting-edge model available on all three major cloud services (i.e. AWS, Google Cloud and Azure).
In addition, NVIDIA and Microsoft will invest up to $10 billion and $5 billion, respectively, in Anthropic to support its continued innovation and development in the field of artificial intelligence.
These initiatives not only enhance the accessibility and utility of Claude's model, but also further solidify Anthropic's position as a leader in artificial intelligence.


Comments are closed.