Pangram verdict · v3.3
We believe this text is mainly human-written, with some AI content.
AI likelihood · overall
HumanArticle text · 755 words · 2 segments analyzed
Something big has happened - a dark horse has emerged among China's top domestic large model developers.A new company, with its first model having only 27B parameters, took second place overall in the CAICT MCP specialized test.Ranked ahead of it is the killer weapon Liang Wenfeng has kept under wraps for nearly a year: DeepSeek-V4-Pro, boasting a parameter count of 1.6 trillion.The difference between the two is a mere 1.3 percentage points.This dark horse is the StartLux-V1.0-27B-Preview, from StartLux (formerly Yuandian Xinghui).Looks a bit unfamiliar, doesn't it? Don't worry, its founder and CEO is an old acquaintance: Chen Danyan.Known as the "godfather" of programmers in the internet era, his achievements are a matter of public record:Shanda Network's co-founder and the head of Lian Shang Network, one of China's earliest programmers to introduce the concept of "shareware"After a decade of retirement, he has made a comeback, this time betting on local models.This has to do with Chen Dawei's recent rare public appearance.At the 18th anniversary reunion of Shanda Innovation Institute, he publicly declared "eight non-consensus views for the AI era," four of which center on local models.Local models will completely destroy the cloud market, catching up with Claude in three years and occupying 80% of the market. The model competition based on parameters is going to be outdated...StartLux is the best representation of his idea.The company's business has not followed the industry trend, instead focusing on the commercialization of small and beautiful local models, making it the world's first truly local model company in the true sense.As the first market-oriented scorecard, StartLux-V1.0-27B-Preview does not rely on the cloud and can run directly on consumer-grade PCs.In other words, this local model, which is nearly 60 times smaller, has Agent capabilities that can match those of a trillion-level cloud-based flagship model.What justification is there for this?Small Model Achieves Big Results, 27B Outperforms 1.6TBefore the answer is revealed, let's take a look at who the comparison is being made to.It's often said that nobody remembers the second place, unless the first is DeepSeek.Moreover, the gap is minimal, making it well worth discussing.The results come from the authoritative institution, China Academy of Information and Communications Technology's trustworthy AI large model benchmark test MCP special item, which sets six types of tasks around real application scenarios:Location navigation, web search, browser automation, financial analysis, code repository management, 3D design.An additional comprehensive assessment will be added, focusing on evaluating Agent's multi-tool collaboration, complex task execution, and interaction in real-world environments.In simple terms, MCP-Universe doesn't evaluate models based on their responses, but rather on whether they can actually get things done. This is also the most fundamental aspect of judging an Agent's quality.The results showed that StartLux-V1.0-27B-Preview had a comprehensive score of 39.25%, ranking second.DeepSeek-V4-Flash-0731, with over 284 billion parameters, and Step-3.7-Flash, at 198 billion, trail DeepSeek-V4-Pro by just 1.3 percentage points.With the same 27B parameter scale, StartLux also surpasses Qwen 3.6 by 5.34 percentage points.It also excelled in individual subjects, ranking first in location navigation, financial analysis, and browser automation, with its other sub-items also ranking high.Let's take a look at two case studies, putting data aside.The first question is about a two-year Microsoft stock investment, with Claude Sonnet 4.6 as the competing topic.Claude's answer is: $47,254, 89.02%.Video link: https://mp.weixin.qq.com/s/365CtdgGYFlEKNDICtoCfgStartLux gave: $47,499.09, 90.00%.It may seem similar, but in the financial industry, a tiny difference can lead to enormous losses.Careful examination of the two models' reasoning processes shows that, due to missing raw data, Claude Sonnet 4.6 misidentified January 8, 2025 as a non-trading day and instead calculated the previous day's closing price.Under the same circumstances, StartLux retrospectively reviews the original data and verifies the market trends around the target date to confirm the accurate closing price before completing the calculation and generating visualization.Ultimately, the conclusion reached by StartLux proved correct, and it was fully verifiable and traceable.The second task was more straightforward: both models were asked to search for flight tickets in a browser at the same time.I'm unable to open a browser or interact with live websites like Google Flights.
I can only process text and don't have real-time browsing capabilities. To find this flight yourself, here's what you'd do: 1. Go to google.com/flights 2. Enter Singapore (SIN) → Beijing 3. Select one-way, set departure date to 5 days from today 4. Filter by "Nonstop" and "Economy" 5. Check the results — note that flights to Beijing may land at either Capital (PEK) or Daxing (PKX). Exclude any Daxing arrivals. 6. Compare prices and pick the cheapest nonstop option landing at PEK.