An intern leads a team, a talent experiment at a startup company.

CN
1 hour ago

"Even a genius's first cry will not be a beautiful song."

Text bySi Wenwen

Edited byZhao Lei

In the competition for AI, both giants and startups are unprecedentedly focused on talent, and they are pursuing efficiency more than ever: the higher the salary, the less patience, and these intelligent young people are best unschooled and untested, able to solve problems as soon as they arrive.

Some young people do not want to be part of an extreme efficiency machine. They yearn for an environment where the company encourages innovation and provides room for exploration, and where personal innovation can also benefit the company.

Liu Zhiyuan, co-founder of Walling Intelligence, believes that as long as a balance point is found, both can coexist. He has taught at Tsinghua University for over ten years, participated in the large model project "Wudao" at the Institute of Artificial Intelligence in 2021, and trained a model with hundreds of billions of parameters, starting Walling Intelligence driven by cutting-edge AI research. A year later, differing from most companies’ pursuit of large models, they chose to take a differentiated route focusing on edge AI. Initially underestimated, but by the time many companies noticed the applications and commercial potential of edge AI, Walling Intelligence was already the fastest-moving entity.

He claims he is not a little genius; the development of Walling Intelligence benefits from a goal-aligned, mutually encouraging, and trustworthy organization. The choice to do edge AI also allows Walling Intelligence to avoid the fiercest battleground, where large manufacturers compete in data, manpower, and computing power, seeing "nurturing talent" as an organizational capability.

Walling Intelligence deliberately retains the atmosphere of Frontier Lab within the company, cultivating talent through a "transfer of knowledge" approach from technical leads to researchers and interns, giving interns certain time to explore the frontiers. This talent program is named "Advancing Four," which means "full speed ahead" in "The Three-Body Problem."

"Part of my motivation for entrepreneurship is to cultivate people at the forefront." Years of teaching have given Liu Zhiyuan a "teacher" quality, always smiling and easy-going, "Outstanding people also need ample opportunities; with enough sunlight, they will surely shine," he emphasized, "Even a genius cannot cry out a beautiful poem at birth."

With limited resources, nurturing talent is even more important

One of the core members in the cutting-edge exploration direction, AI4AI, at Walling Intelligence is an intern named Li Qi. He looks very young, graduating with a doctorate next year, and only joined Walling Intelligence last year. The team he leads has dozens of people, making it the largest group in the smart evolution laboratory dedicated to cutting-edge directions. Close to half are full-time researchers, having been at Walling Intelligence longer and possessing more experience than he does.

He certainly meets the standards of an excellent intern: winning a national physics competition gold medal, directly admitted to Tsinghua's Yao Class for undergraduate, and pursuing a doctorate in automation at Tsinghua. However, large companies might see this as somewhat risky: after all, interns are still inexperienced and are considered somewhat unstable. They prefer to manage individuals with some experience or those who have spent time in well-known companies.

"AI is a field that requires exploration and trial and error," a Walling Intelligence management member responded, "Who starts off as an expert?"

In the first six months of joining Walling Intelligence, Li Qi explored several directions in the AI Infra team, with AI4AI being the most interesting to him—using AI to research, train, evaluate, and improve AI, allowing for self-iteration, potentially breaking the limitations imposed on AI capabilities by human data and computing power supply, accelerating AI R&D at an exponential speed. However, he felt the path was not clear enough and the timing was not yet ripe until February 2026 when Claude Opus 4.6 was released, demonstrating AI's ability to independently complete complex R&D tasks.

"It seems this story can progress!" said Li Yuxuan, the head of the Smart Evolution Laboratory, to Li Qi, "Let's give it a try." Li Yuxuan is also only 31 years old and understands the need to encourage young people's ambition.

They started testing from the Infra level, with several people spending a month validating that AI could assist in R&D at the Infra level, such as writing code, debugging, and cleaning data. Walling Intelligence decided to give Li Qi more time, team members, and resources. Li Qi and several researchers began to fully immerse themselves in AI4AI, and two months later, under the leadership of interns, they created a large model pre-training framework called ForgeTrain, entirely written by AI, with no human involvement. On Nvidia H100, the training speed was about 10% faster than the mainstream framework Megatron used by most large companies and startups.

After the project was released, the AI Infra department was upgraded to the Smart Evolution Laboratory, with Li Qi naturally becoming the head of the AI4AI direction. Liu Zhiyuan predicts that in six months to a year, AI4AI will enter the engineering phase, practically affecting model production, with only a few companies able to seize the opportunity.

This appointment went through little discussion, "Capable, everyone can see," and no specific KPIs were set.

One of the goals of the "Advancing Four" talent program is to transform this "atmosphere" into a more solid and explainable nurturing system. In 2025, the explosion of DeepSeek made the industry aware that a few people could create enormous returns, and both large companies and startups began recruiting talent at high salaries. Han Xu noted that it has become challenging to recruit for almost all positions at Walling Intelligence; many came for interviews, but when offers were made, large companies would also increase their offers. When Walling Intelligence was just established, it was unrealistic to compete with large companies on salary, "If we don't have more money than others, should we clash head-on? If we rely solely on high salaries, the recruited individuals may not even strongly resonate with us."

Some large companies prefer to screen rather than nurture talent, pursuing extreme efficiency, hiring individuals who can work immediately and rapidly cutting redundant manpower when necessary, willing to pay higher salaries for this. This might be an opportunity for Walling Intelligence, with chief researcher Han Xu summarizing it as "approaching the talent shortage problem from the perspective of nurturing people."

He discussed with Liu Zhiyuan and other leaders, deciding to start with interns and establish the "Advancing Four" talent program. If they believe edge AI has long-term potential, requiring five, ten, or even several decades, they need to patiently build a talent pipeline and enhance team capacity through practical training.

"We are one of the best teams in China at cultivating AI talent; why not utilize this advantage? We can help interns become excellent at Walling Intelligence." Counting from when he worked as an assistant researcher, Liu Zhiyuan has taught at Tsinghua for thirteen years, being part of the earliest batch of NLP researchers and nurturing numerous students, "We do not require interns to have experience or immediately be able to work, nor do we expect them to necessarily stay at Walling Intelligence in the future. To be honest, if we can cultivate truly outstanding individuals, that also contributes to the country. Besides, in the next few decades, there will always be other opportunities for collaboration."

They made it clear that the "Advancing Four" talent program aims to make interns feel recognized, with some of the most outstanding interns receiving incentives like options before graduation; involving interns in core projects, allowing them to independently handle projects rather than just performing basic, repetitive tasks; and providing interns with additional resource support to get R&D results applied to company products faster.

Before interviews, Li Qi had not conducted research related to AI4AI; he discussed more about understanding AI, how to use AI to solve problems, and the logic and autonomy behind problem-solving with the interviewers. Another intern in the "Advancing Four" talent program had an undergraduate degree from a lesser-known institution and met the Walling Intelligence team while participating in a supercomputing competition, connecting over similar topics; he wrote an email to Liu Zhiyuan expressing his keen interest in large models.

The response came quickly. After a technical interview, he joined Walling Intelligence when there were only nine people in the company, starting work on the text large model infrastructure.

After working for more than two years on the text large model infrastructure, this "Advancing Four" intern felt that his growth space had diminished and much of his work had become repetitive. He conducted some research and directly told Han Xu he wanted to shift from text to Omni (all-modal). Han Xu quickly helped him contact several senior researchers to see if he was truly interested. If confirmed, "No problem, let’s get started," he paraphrased Han Xu's answer, admitting he hadn't overthought it because he trusted Han Xu would accept his ideas and match them to his interests. Sometimes he felt more like a student at Walling Intelligence than an employee.

"The company nurtured people, and they also helped the company complete tasks. It's a win-win for everyone," said Han Xu. Some interns want to go to large companies or switch to academia; he helps by recommending them and providing resources, "If an intern can achieve a high position in another company, that's at least proof that our cultivation was successful."

This is partly related to the positioning of the "Advancing Four" talent program, as interns are the most flexible group in a commercial company, still in school with ample choices, relatively free from KPI pressure unlike full-time researchers, who do not need to be responsible for commercialization. Those eligible to join are also the few interns with great potential; the company hopes they can explore and innovate in a supportive environment. On the other hand, DeepSeek relied on many recent graduates and interns to produce world-class outcomes, showing AI companies the immense benefits of concentrated talent, leading everyone to want to enhance their talent attraction and shape better employer branding.

At the end of May, the ForgeTrain project led by Li Qi was announced and open-sourced, and the interns involved received invitations from several major talent programs in large companies. Walling Intelligence also received many internship resumes, most of which mentioned the ForgeTrain project—"If willing to let interns undertake tasks where the correctness is uncertain, success timing unknown but innovative, it represents some ideals," Li Qi said, those wanting innovation would be attracted to organizations willing to encourage innovative talents.

Innovate like Frontier Lab

When Walling Intelligence faced its most challenging decision since its establishment, Liu Zhiyuan brought over 20 people from the company together to read "On Protracted War." At the end of 2022, the release of GPT-3.5 let investors and entrepreneurs see the potential of large model companies, driving the valuations of the AI "Six Dragons" high; even though the Walling Intelligence team had produced foundational models with hundreds of billions of parameters, market enthusiasm was not high.

Two paths lay ahead: continue making larger models, such as a 140B model, investing over a hundred million to strive to become the next OpenAI, or abandon the pursuit of super-large models to turn towards edge AI?

Liu Zhiyuan learned from "On Protracted War" that as an underdog facing a stronger enemy, one should proactively seek opportunities, forming absolute advantages on localized battlefields—hence when startups compete with large companies on who can burn more resources, it is an unequal fight. At that time, he judged that in six months, at least seven or eight manufacturers would be able to train models at GPT-4 levels, and a price war would follow. Walling Intelligence's advantage does not lie in competing for resources.

"Within the range of large companies, it's certain death," in that office, everyone present shared their thoughts until they reached a consensus to pursue differentiated edge AI. In the second half of 2024, they proposed the Densing Law of LLM: the number of parameters required for large models to reach a certain level of intelligence halves every 3.5 months, later published in a sub-journal of Nature.

Walling Intelligence's core products, such as the MiniCPM edge large model, are all based on this technical judgment. After 2024, they gradually received recognition from investors and the market, "because everyone began to believe that we could drive innovation through technical advancements."

"Organizational density is the foundation of innovation. The difference is whether, when results are poor, we think of ideas together, or blame each other. If everyone is grouped in combat, with the data, model, and Infra teams sharing common goals, they realize that each person is crucial to the final outcome." Liu Zhiyuan examples that when a company develops a large model, participants often exceed 2,000, while DeepSeek achieved more leading results with just over 100 people. Liu Zhiyuan and Han Xu discussed the organizational traits of DeepSeek, concluding that the organization is efficient, has low internal friction, and is united in purpose and beliefs.

He explains that the core of AI companies is to systematically train leading models. Single points such as model architecture and reinforcement learning require in-depth research and innovative methods to accelerate, serving as excellent frontier topics. If graduate students do not participate in specific production but instead publish papers in silence, they lack training resources and become increasingly distant from the frontline; with thousands of papers published each year, very few are read, and they hold little practical significance.

If the single point of a company's technical system does not continue to innovate, it will also lose competitive edge; without achieving a closed loop of industry-university-research, front-line AI is hard to achieve, "If people are well-cultivated and innovative, Walling Intelligence will develop better. Conversely, it is also true."

They position Walling Intelligence as a hybrid of Frontier Lab and a commercial company. For instance, interns usually stay within a company for half a year to a year; if the team undergoes major adjustments, research topics may be difficult to continue. Liu Zhiyuan states they will strive to create a long-term stable environment, allowing interns to focus on research.

After spending years at Walling Intelligence, an intern from the "Advancing Four" talent program began to take responsibility for parts of the R&D of multimodal models, gaining permission to hire and lead others. He said the company did not impose limits; however, after discussing with the leads, he established a rule: first clarify what new colleagues can clearly do in the recent 3 to 6 months and what they might do in the next year or two—the core is ensuring everyone has enough space to thrive.

Competition among model manufacturers is intense, with version update cycles continuously compressed, leaving many companies unable to consider these. Development of the company and personal growth do not always align. However, this intern joined Walling Intelligence when it was just starting, participating in pre-training, post-training, model releases, and open-source maintenance for multimodal models; the technical scope he could access is broad enough that he describes the team atmosphere as ensuring 80% deterministic tasks with clear returns while exploring 20% uncertainty, finding advantages through non-consensus solutions.

"If new colleagues can only do basic tasks with narrow imaginative space, it is detrimental to them and the team's atmosphere. Can efficiency really be high?" He hopes the team still engages in thorough discussion and refrains from hiring excessively.

One intern, at first not considering it, engaged in a research topic for his thesis, studying how to convert a 1B dense model into a sparse model; initially, no one knew how likely success would be, yet the company provided plenty of computing support. Once validated on the model, he shared the good news with Liu Zhiyuan; a month later, his experimental project was applied to the company's products.

This became his most impactful project at Walling Intelligence. "Very quickly, really fast!" The company might not support interns' experiments, and even if successful, it could take months to validate. This could be the norm. After the results were published, he was included in the "Advancing Four" talent plan and took part in more core work at Walling Intelligence—indeed, his innovations benefited Walling Intelligence.

The characteristic of AI companies is to start from cutting-edge issues or unverified research directions, not only productizing existing technologies but pushing past "impossible" to "possible," and then converting "possible" into real value. This accelerates the conversion between innovation and commercial benefits, allowing industry-university-research to close the loop in shorter cycles. However, whether the innovative atmosphere can be maintained tests the founders' will, team values, and long-term patience and determination.

From "Thoroughbred" to cultivating "Talent Scout"

Han Xu does not feel there is much difference between working at Walling Intelligence and teaching students at Tsinghua. "Both hope to cultivate a group of people, to fight together and accomplish tasks." He pondered for a moment, "The difference may be that a single hammer cannot drive too many nails; in a company, there are more hammers to drive."

He has always wanted to be a teacher, educating others and achieving widespread recognition. This has been somewhat influenced by his mentor Liu Zhiyuan. After more than four years of entrepreneurship, growing from a company with fewer than ten people to over 400 and preparing to go public, Liu Zhiyuan believes the mission of Walling Intelligence is to transform cutting-edge technology into business success, but his personal primary mission remains nurturing talent.

Many employees at Walling Intelligence habitually refer to Liu Zhiyuan as "Director Liu," not only because he has taught at Tsinghua University for over a decade but also because he retains the caring demeanor of a counselor toward students. Intern Li Qi occasionally encounters Liu Zhiyuan at the company; they rarely talk about technical details, instead discussing internship experiences and future plans.

He considers himself a beneficiary of being "cultivated": when he took the entrance exam from Shandong to Tsinghua, he had no research papers, no competition experience, and his grades were not outstanding; he was just "nobody," initially thinking that graduating smoothly was already quite good. However, with the help of his mentor Sun Maosong, he gradually worked his way through, obtaining an excellent doctoral thesis and engaging in research, becoming a "thoroughbred" cultivated by a talent scout. During his master's studies, he saw some undergraduates struggling initially, but with patient guidance, they transformed significantly within two to three years.

"So, thoroughbreds are often found, but talent scouts are rare." As a mentor and entrepreneur, he sometimes notices young individuals initially underperforming and thinks that they are certainly smarter than he was back then, "So I must work hard to ensure they shine too." He hopes to be a "talent scout" alongside the technical leads at Walling Intelligence, cultivating interns into "thoroughbreds," and then those "thoroughbreds" can continue to be the next generation of "talent scouts."

This goodwill is gradually transmitted downwards at Walling Intelligence, naturally forming a certain atmosphere.

Han Xu agrees with the perspective of Google DeepMind research scientist Yao Shunyu, stating that the era of individual heroism in AI has passed; the most important trait in a collective is reliability and accountability, "As a tech person, if you idolize little geniuses, then there's something wrong." He values whether candidates have a long-term consideration for AI and if they fit well with the team.

Interns sometimes consult Han Xu for rental experiences; for some academic research topics not closely related to the company, he is also willing to give advice and introduce relevant researchers. He now spends two-thirds of his time on nurturing-related matters, "Many times, first thing in the morning, when I open WeChat on my phone, it’s all 'bad news'." But he feels it's a good thing, "When people truly have issues, they come to me."

"You see how others help you, and then you are willing to help others." Many interns state that they find it hard to explain why institutional constraints are unnecessary, "Maybe it’s simply empathizing with each other."

One intern received invitations from several large companies, stating one reason for staying at Walling Intelligence was finding career role models here who are technically capable and willing to mentor newcomers—"Having resources in the company is one thing, but resources do not automatically transform into capabilities; the key is knowledge transfer and support." Initially joining the company, he observed how the technical lead conducted experiments, analyzed errors when experiments failed, and summarized insights when successful; "It was like reinforcement learning and supervised learning," from figuring out how to solve a bug to establishing a set of experimental methodologies, now, with experience, he has become the one mentoring new interns.

Many interns share this mindset. An intern in the "Advancing Four" talent program hopes his new intern can independently handle an RL Infra open-source project, and he helps set goals and processes while allowing the other to take practical steps. This project does not directly relate to Walling Intelligence's business, but he assured during the interview that he would support the exploration since he was treated this way.

Recently, Li Qi discovered that some newly recruited interns were not adapting well. He reflected with the business leads that since they passed the interview, it signifies potential. Why wasn’t that potential being realized? After filtering, they believed that the guidance for newcomers was inadequate; there was no systematization regarding objectives, which documents to review, what coding tools to use, etc. He and several technical leads spent time preparing a detailed document comprising both technical concepts and practical details like office layouts.

This type of help is also recognized within the organizational system. Walling Intelligence implements employee peer evaluation, with some colleagues clarifying in their assessments that information communication and helping others also count as contributions. A monthly TGIF all-hands meeting is held, where Liu Zhiyuan and others discuss recent progress and future planning, alongside a segment called "Culture of Partnership": who helped you, allowing for public gratitude.

This is related to Liu Zhiyuan's hiring philosophy. He does not believe that just recruiting a few well-known researchers from large companies can resolve issues; he values nurturing from within, retaining recent graduates and interns to grow into technical backbones. He previously shared a thought on a social platform: if you help someone acquire knowledge, build confidence, and find a goal that brings them genuine happiness, then once their potential is unleashed, their power will be immense.

What if well-trained interns want to leave? An intern recently discussed this issue with Liu Zhiyuan, who advised him to accumulate a few years of experience and material conditions at Walling Intelligence first, regardless of whether he intended to start a business, join another company, or pursue academia, then choose the most suitable path and aim for top-notch performance.

Maintaining morale through continuous victories

A researcher at Walling Intelligence gradually felt that efficiency was declining. Previously, he was just an aisle away from the technical leads with the company totaling a few dozen people, and any issues were quickly communicated. Now, the company has expanded to over 400 people, with workspaces spread across several floors, longer meeting durations, and sometimes difficulty booking meeting rooms. Some also sensed shifts in the company's business focus. Liu Zhiyuan used to emphasize that models needed to be SOTA, while this year he discussed download volumes more, incorporating it into some researchers' KPIs.

As the company grows, scales, and faces heavier commercialization tasks, organizational efficiency and values will progressively encounter challenges in understanding and communication. Liu Zhiyuan has heard some discussions, "We must remain pragmatic. Model companies all talk about innovation; how to evaluate effectiveness? Then see who better meets user needs." Mid last year, the team was somewhat disheartened due to text model performance not meeting expectations. Upon tracing through every aspect, it was found that data, evaluation, and other infrastructures were not solid enough. They spent over half a year addressing shortcomings; resolving engineering issues seemingly lacked the thrill of technological innovation, but business organization could not be avoided.

The R&D personnel across several laboratories at Walling Intelligence typically comprise thirty to forty members, marking the boundary Liu Zhiyuan considers efficient communication for a technical lead. Recently, he discussed with Han Xu and other leaders how to maintain high organizational efficiency in an AI Native manner?

They did not want to use traditional daily or weekly reports, as this might just pile on additional burdens without aiding efficiency, "Resolving leadership's insecurities." Han Xu learned from some non-commercial organizations about controlling personnel scales with technical means, reducing management costs while increasing combat effectiveness. Recently, Walling Intelligence began promoting the use of AI to simplify processes, encouraging researchers to develop various Agents for running experiments and basic analyses. The research outcomes of AI4AI, ForgeTrain, have also entered production processes, starting to optimize model training, deployment, and other areas, taking on some repetitive tasks.

Within each section of Walling Intelligence, exploration is also ongoing. The multimodal model R&D team has roughly doubled in size over the past year. A technical lead discovered that meeting times were becoming longer and numbers growing; many colleagues attended simply to synchronize information. He tried adjusting the meeting structure to a "fork" format: people with similar issues meet in one session, sending a representative to larger forums for information exchange. This has reduced some meetings from dozens of participants to only ten, needing to attend just one meeting per week.

Another solution is to have the technical leads bolster management. As outstanding young individuals increase, the company began encountering new issues: when several talented interns or researchers form a group, every member is skilled but struggles to collaborate effectively.

“You can’t just vaguely shout slogans like ‘everyone needs to unite’; no one will take that in.” He summarized experiences in company management, “Doing models in a company combining production and research is like launching a rocket; there needs to be precise milestone estimates, steady progress while ensuring engineering while conducting forefront explorations full of uncertainties.” He has confidence in the tradition of altruism within the organization; it is just "very complex and requires learning."

Recently, he is preparing to organize leadership courses for the technical leads to learn about management, from how one intern can motivate two or three colleagues, to how an organization can evolve, ensuring every direction's lead is motivated and capable of effectively guiding their teams.

Having been in entrepreneurship for over four years, he has realized that each stage presents challenges: initially, the technical concept was hard for investors to recognize; now it is about maintaining technical advantages while converting them into sustainable scalable income. The commercialization of Walling Intelligence currently mainly occurs in the automotive and mobile phone industries, also making it a preferred choice for certain large model manufacturers transitioning from cloud to edge, which tests not only technical capabilities but also engineering adaptation, compliance experience, industrial resources, and partner willingness.

Competition is becoming fierce. His vision for Walling Intelligence is to fully realize the state of general AI edge models in the next three years, achieving sensory, computational, and action capabilities and enabling soft and hard collaboration.

In this process, the team will also expand accordingly. Liu Zhiyuan and Han Xu believe that fundamentally, the best way to maintain organizational combat effectiveness is through continuous victories.

Some companies can scale their computing power and data following the Scaling Law, solving issues through engineering methods, while others have advantages in industrial ecosystems and resources, but the most crucial development for startup companies like Walling Intelligence comes from innovation. The discoveries and applications of the Densing Law in 2023 and 2024 have shown investors that in addition to pursuing super-large models, optimizing edge models through algorithms, data, and architecture can reduce inference costs while ensuring performance, enabling large-scale application in devices like mobile phones and automobiles; explorations ahead of time in AI4AI's ForgeTrain project confirm AI's capacity in model production, and perhaps users can adopt more efficient training frameworks depending on differing models, chips, and training tasks, exploring outside the Nvidia CUDA software ecosystem.

In the past year, an intern in the "Advancing Four" talent program worked with the team to develop the Omni-Flow streaming all-modal framework, improving the interaction pattern between users and AI from half-duplex (question-and-answer) to full duplex (parallel communication, actively responding or interjecting at any time). "Training super-large models using thousands or tens of thousands of cards at a large company sounds impressive, but if the training method and objectives are set by referencing another large company, it is not exciting either." He enjoys exploring lesser-validated technical routes.

This aligns with Liu Zhiyuan's initial expectations of entrepreneurship, "Everyone working together to achieve what seems impossible in the AI era, while also standing at the forefront to cultivate talent— isn’t that the most wonderful thing?"

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink