This guy is queuing in Venice to walk the red carpet while winning two AI image awards.
Many people may have heard his name.
DiDi_OK.

A few days ago, DiDi_OK released his latest work "Candy," made with Seedance 2.5, which has a Chinese title "糖果". After seeing it, I was amazed.
I directly forwarded it, recommending everyone to study it frame by frame.

Although in the AI film industry, the name DiDi_OK has almost become a household name.
But many friends may still not know this name, but that’s okay, you’ve probably seen his previous works.

For example, "Arrow," "Website," "Garbage Station," and many more, while the best data was from this year's viral "Brand."
This film had 20 million views just on Bilibili, and close to a hundred million across the internet.

I still remember the night I first saw "Brand." After looping it six or seven times, the thought that arose in my mind was:
“Clearly we are all human, why do I feel like a drooling fool in front of DiDi_OK? What kind of mind and aesthetic does this person have?”
It’s not an exaggeration to say this is a genuine sentiment; since then, DiDi_OK has become the undisputed number one in my heart for AI imagery.
And this time, "Candy" is DiDi_OK's peak work in this parallel universe series, bringing technology and creativity to a high climax.
If you haven't seen it, you should watch it once; trust me, you will definitely be amazed.
Unlike many greasy and cliché AI works seen today, "Candy" still retains a strong Di flavor.
There are some high-dimensional rules presented, a sense of SCP, and some echoes of "Black Mirror."
There’s even a shadow of Liu Cixin.
Of course, there are also the strong narrative rhythm, camera language, and his aesthetic.
So after the release of "Candy," I felt it was an opportunity.
I wanted to chat with him, to do a small interview and discuss everything about him.
As he is overseas and has been at the Venice Film Festival receiving awards, we had some time difference; after two days of trying, we finally scheduled a meeting.
On the day of the interview, it was 10:30 AM in Beijing.
He was in Montenegro, where it was 4:30 AM; he got up early because he had to head to the mountains around 6 AM...
We had a great chat for nearly two hours, asking many questions I wanted to know. After our conversation, I thought I must write down these things and share them with everyone.
Now, let’s begin.
I. He Welcomes the Age of AI
DiDi_OK is a very interesting person.
He has been living in London long-term and working at WPP Production.
WPP is one of the largest advertising and marketing communications groups in the world; the well-known Ogilvy is a subsidiary of WPP...
And WPP Production is the core production team under WPP.
So before the arrival of AI, DiDi_OK had already been working in animation, 3D, and gaming.
His undergraduate and master’s degrees are in animation.
Later, he joined WPP, starting with animation as well.
His story about going to the UK is also particularly interesting.
At that time, he was earning money by running a studio with friends in China, mainly shooting ads.
As he was doing this, he suddenly felt it was boring...
One afternoon, they were even shooting ads outside.
He suddenly asked his friend: Why don’t we go abroad?
His friend asked: Where to?
He said: How about the UK?
His reason for going to the UK was rather ridiculous,
because after shooting projects, they would stay up to watch "Film Hurricane," and Tim was also DiDi_OK's cyber idol.
Moreover, Tim used to be in the UK...
So why not just go to the UK = =
Then he really went home and studied English for half a year, applying to the University of the Arts London.
His mother initially didn’t help him apply because she thought he would definitely not get in.
In the end, he applied secretly and actually got in...
Years later.
This guy who used to watch "Film Hurricane" at midnight after shooting commercials was invited by Tim to sit in the AI course of "Film Hurricane" and share his creative insights.

You see, the world is always a wonderful circle; life sometimes truly resembles a somewhat cliché screenplay.
Before AI, he had already started creating.
In 2021, he produced a work called "Moth" at the University of the Arts London.

And even then, it had a strong Di flavor.
A closed world shrouded in fog, with a massive tower at its center.
Everyone believed that it must not be touched, but the protagonist, like a moth to a flame, had to go up.
Even if the truth is disappointing, even if it could lead to death.
They still wanted to see the true nature of the world.
In his mind, there were too many ideas about technology, the unknown world, and civilization.
But in the past, what had held him back was that era's productivity.
The workflow for traditional animation was too lengthy, making it nearly impossible for one person to create a film independently.
The more processes there were, the more people involved, and with each additional person, the ideas in your mind would lose a bit in translation.
Then, experiencing modeling, rigging, animation, post-production, and so on would lead to further losses.
In the end, that idea in your mind, the one that excited you and gave you goosebumps, when it appears on the final screen, might not even retain half of its essence.
So, over the years, due to these losses, he kept researching something he refers to as:
Strength.
How to convey the creativity and feelings in his mind to the audience with minimal loss.
And finally, he welcomed his era.
AI has arrived.
II. "Candy": Bad People Will Explode
If you ask me to introduce "Candy."
I might carefully explain for half a day:
It’s a science fiction story about first-level civilization, alien rules, violence, order, humanity, and technological development.
However, when I asked DiDi_OK to explain, he thought for a moment and summed it up in about ten words:
“Bad people will explode, while good people will be fine.”
Extremely straightforward, yet precise.
At the beginning of this story, humanity's energy usage capacity has reached the threshold of a first-level civilization.
Then, a moon of Saturn, called "Pan," comes close to Earth.
Subsequently, an absurd event starts occurring worldwide.
As soon as someone prepares to harm another person...
Bang, they explode.
This explosion doesn’t turn into blood and flesh, but instead becomes:
A pile of candy.

At first, everyone was confused.
Alien? New type of weapon? God?
Gradually, it was discovered that this rule specifically targets:
Violence.
The underlying judgment criteria set by DiDi_OK have a strong Chinese flavor.
It’s just four characters: "Intent to harm."
After the film was released, I watched it four times. At that time, I asked him about a topic that felt very contradictory. It seems that this is clearly an anti-violence story.
But the greatest pleasure for the audience is watching the bad guys being taken down by a stronger force.
Isn’t this inherently violent?
DiDi_OK laughed at that time, saying you’ve grasped the essence, and that’s actually what he intended.
The film later sends repeat offenders to the moon.

It seems too good, civilization has become peaceful.
But behind this, there remains a form of collective violence.
So he even spent a long time designing a female spokesperson in the film to represent public opinion.
He wanted her to possess a particularly easy-to-accept and trustworthy quality for the public.
Somewhat akin to Cheng Xin from "The Three-Body Problem."
Because this matter itself carries a certain danger of "everyone thinks this must be right."
DiDi himself does not stand on the absolute anti-violence side.
He believes there’s often a cruel question in reality:
Good people are often those who get bullied.
Hence, from Western knights in the Middle Ages to Chinese chivalrous heroes.
Humans have been fantasizing about one thing for thousands of years:
Justice is best accompanied by power.
At this time, the core of "Candy" became different.
An alien moon has the final say on what is evil, and all humanity, due to fear of punishment, enters civilization, then collectively agrees to send noncompliant individuals to the moon.
Is this considered civilization? Or order? Or tyranny? Or a world that is slightly better than ours now?
DiDi_OK has no answer.
This question belongs to each viewer.
III. The True Progress of AI
Before the interview, I re-watched DiDi_OK's films, including "Arrow," "Brand," "Website," "Garbage Station," etc.
One of the most obvious changes is that "Candy," made using Seedance 2.5, finally dares to slow down.
For example, this textbook-level long shot.

To be honest, in the past, DiDi_OK's films used to have frequent quick cuts, like "Arrow" and "Brand," often changing scenes every five seconds.
The reason is very simple; people in the AI image industry might guess this practical reason.
You don’t dare to slow down because the model’s capability simply doesn’t allow it.
The past model capabilities could not support the storytelling of long shots where details remain unchanged.
Thus, one could only keep using new visual stimuli to cover up problems that the model couldn’t reveal in time.
However, in "Candy," for the first time, there are large amounts of quiet narration.
For example, this shot.

In the laboratory, an older official walks forward.
A young researcher follows closely, with a very obsequious attitude, eager to be noticed.
But the official remains cold until he suddenly acknowledges him with a remark.
The young man’s body immediately bends slightly, and his eyes light up.
At this moment, you can see:
“Hey, he heard me speak.”
But DiDi said this is where he feels the model truly broke through.
Performance.
He always felt that AI performances couldn’t pass muster; the details of those performances weren’t enough, so there was always a lack of these narratives.
Now, in his eyes, the gap with Seedance 2.5 has officially been crossed.
The girl at the beginning eating candy is also significant.

He particularly cares about that small pimple on the girl's face.
I didn’t even notice it the first time I watched.
But for DiDi_OK.
This is an extremely important information anchor.
Because this minuscule detail will make the audience feel at first glance that this is a real person.
When an actor's face can provide information, the director dares to keep the camera there.
Moreover, in this close-up of the fingernail, the texture of the skin and nail.

Also, in this close-up, the character presents highly precise microscopic details of the skin, including pores, fine lines, sebaceous gloss, and changes with light; at the same time, subtle performances such as eye expressions, breathing, muscle tension, and micro-expressions all possess credibility.

And this point, DiDi_OK says, this level of detail and refinement can only be achieved by Seedance 2.5.
He believes this model has saved him at least 60% of the narrative costs.
I particularly like this term.
Everyone talks every day about how AI reduces production costs.
What he talks about is narrative costs.
I think the real change in narrative has never been just higher resolution or cheaper prices.
It’s that in the past, you could only tell your story around the model's flaws.
But now, the model allows us to tell stories according to the ideas in our minds.
That is the true progress.
IV. Everything Changed by Seedance 2.5
We talked a lot about the Seedance 2.5 model.
DiDi_OK said that if someone told him now,
that he would never see a new video model in his lifetime,
only could use Seedance 2.5 forever.
He would feel that's enough.
Because frankly speaking, from my perspective, the reaction online when Seedance 2.0 was released was far greater than the response this time for 2.5.
But at that time, DiDi_OK felt that 2.0 was okay, indeed shocking, but it didn’t feel like a significant revolution.
Until this time, with Seedance 2.5.
Many people might first think, huh? Isn't it still an AI video? The difference with 2.0 isn’t that great.
But DiDi_OK feels that this is actually a genuine cross-era upgrade.
I asked him: Where is the crossover?
He mentioned two words.
Highly customizable.
High-precision motion laws.
These two terms sound particularly technical.
Especially “motion laws,” which just doesn’t sound as attractive as the terms “cinematic feel,” “1080P,” or “one-shot.
But the audience's eyes are particularly honest.
For example, when a car turns, the body should lean, the suspension should compress, and braking and acceleration should create different pitches.
A piece of candy explodes out from a human body; it should have mass, initial speed, bounce when it hits something, and there will be collisions between hundreds of candies; depending on the distance from the camera, the degree of blurriness will also differ.
The so-called motion laws are essentially: does this world conform to the physical laws we learn from our understanding of the universe.
So the most core "candy explosion" in "Candy" is actually a limit test for AI capabilities.

DiDi_OK said that the visuals produced now by Seedance 2.5 can endure slow motion and pause.
He even intentionally designed shots to let viewers pause and look.

For example, in this police station scene, there are many people in the shot, and DiDi_OK has created a separate storyline for a black character in the background.
First breaking free, then pushing a table, picking up a TV, preparing to attack the police, then the lights flicker, and everyone's actions pause for a moment, while specific people simultaneously explode into candy.
DiDi_OK leverages the model’s capabilities to simultaneously direct the actions of over ten individuals, each with their storyline.
This is what DiDi refers to as the second characteristic: customization.
Because in the past, making a monster or something with AI wasn’t particularly hard.
What’s difficult is often very mundane, like "have the third person on the left look back first. Then pick up the TV. The police on the right continue their current action. After the light flashes, the second person explodes."
Once AI truly enters industrial production, the most needed capability is precisely this kind of customization.
DiDi_OK says, from a visual perspective, he feels there is almost nothing that he wishes to realize that cannot now be achieved.
For example, a bunch of candies forming Matisse's "Dance."

Time-lapse animation in the miniature landscape.

And so on and so forth.
Due to the enhanced model capabilities, motion laws and customization have become extremely strong, which has also brought another change: the previous workflows have somewhat failed to keep up with the versions.

For instance, in the past, to create a shot, it might require several reference images to help the model understand.
Therefore, he said at that time, he had a clear priority:
Reference images > Prompt > Model selection.
Even to enable the model to do animations and camera movements that its own capabilities couldn’t achieve, he would extensively use Blender to build models for reference.
For example, motion tracking between characters and surfaces was previously inaccurate. So one had to track each word one by one, all of which would need manual post-production repairs.
Basically, it was a skilled human expertise.
Essentially simple, it's due to the model's capabilities falling short, and DiDi_OK didn't want to lower the quality of his works, so he would use many other methods to make up for it.
Then came "Candy."
He said this logic has clearly started to reverse.
The importance of reference images has significantly decreased.
The importance of prompts has started to skyrocket.
Because the model's comprehension ability has strengthened.
Many animations that previously required keyframes to be completed can now be directly articulated.
All traditional assistance has become useless, and a lot of the additional post-production that used to take two weeks has disappeared.
DiDi_OK said this is his first nearly 100% pure AI film.
In fact, before using Seedance 2.5, he had already made 3 minutes of "Candy" using 2.0, nearly completing 40%. After using 2.5, he resolutely scrapped it all and started over.
This is also one of the frustrating aspects of the AI era.
The skills and workflows you meticulously honed over the half-year might just be completely swallowed by a model release.
DiDi_OK said those are all just technicalities; he is increasingly uninterested in them. What truly determines your ceiling is still your personal ability, and the model’s comprehension ability itself.
When the model starts to consume workflows.
What remains for the creator is what truly belongs to the creator.
V. What is Creation Ultimately?
I asked DiDi_OK: “Do you think AI can really be considered creation? Many people say, isn’t it just drawing cards? Where is the creation?”
He gave me a particularly good example.
In 2023, he participated in an F1 racing promotional video, combining live action and virtual production.
At that time, many shots might need to be retaken over 500 times.
Just the venue cost for a day was £100,000.
The car would drift by, the tires would squeal, and smoke would rise, but the director would think the smoke looked bad and call for another take.
Every setup would have to reset everything, and then they would go again.
Like this, continuously rolling and selecting.
Until the director thought that the smoke from the 397th take was the best, and used that shot.
So, DiDi_OK said, creation has never been about how many times you generate; it’s not about whether it’s AI or not, because essentially, traditional filmmaking also follows a logic of trial and error.
True creation is that "no."
Why would this director want the 397th take? Why weren't the previous 396 good enough?
This "no" represents aesthetics, judgment, experience, and that is what true creation is.
It’s just that under the support of AI, the speed of drawing cards has become increasingly fast, and the quality higher, far surpassing what traditional filmmaking could ever achieve.
But paradoxically, after the model becomes stronger.
DiDi feels no more relief at all; in fact, it has become even more painful.
A past 8-second shot might take about two to three days, because he knew how the model worked, its limits were set, and he could comfort himself that it wasn’t his fault.
But now, with a model like Seedance 2.5, you can get an incredibly powerful result by the third take.
That means it’s possible.
At this point, all the pressure is shifted to DiDi_OK.
Why can’t I orchestrate it?
Why didn’t I think of that angle?
Why isn’t it good enough yet?
Previously, a shot would take two to three days; now he may spend an entire week refining one.
He said:
“The model can clearly do it, but I can’t produce it, so obviously that’s my own problem; I’m still too inexperienced.”
The AI era is also a cruel era.
As the cost of expression decreases,
mediocrity will have fewer excuses.
In the end.
I asked him a significant question: “You’ve filmed so much about civilization, rules, society, and humanity, do you still believe in humanity?”
He said: “If we talk about humanity itself, then I must be an extreme pessimist.”
He said he sometimes falls into a severe existential crisis.
When AI can help you accomplish more and more tasks,
what is left of you?
The newcomers in his company often ask:
Teacher, what should we learn in the future?
He also doesn’t know.
Then he said what I found to be the most moving sentence throughout the entire interview:
“When AI can help me do everything, am I really an interesting person?”
There are no answers, just like his works.
Written at the End
After the interview ended, I watched "Candy" again.
I suddenly found this film easier to understand.
DiDi_OK clearly says he is pessimistic about humanity, yet his works always have this peculiar underlying tone:
Hope.
I suddenly felt that in him, I saw the shadow of a character from "The Three-Body Problem," namely Wei De.
He might not usually be very willing to interact with people, and sometimes have some pessimistic judgments about humanity as a species, but in the end, like Wei De,
he does not wish for humanity to lose.
I particularly like this contradiction, as many great creators are often like this.
They create.
Rarely because they think the world is already good enough; rather, they feel it is still not good enough.
That’s why they create another world.
One that allows many people to feel the beauty with them.
A wonderful world.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。