What If Stallone & Schwarzenegger Made CYBER SLAYER (1995)?

Hey all!

This started as a way to mark my 15th anniversary working in video game cinematics. I thought it would be fun to make a completely ridiculous, fictionalized version of how I got into the industry, presented as the trailer for a big 1995 action movie.

It started fairly small, but the technology kept improving while I was working on it. Every time a new model came out, I started thinking, “Maybe I can actually make that shot now.” Eventually it became much more ambitious than I had originally planned.

This subreddit was a huge help throughout the process. I found a lot of technical solutions, new models and inspiration here, so I wanted to share the finished trailer and explain some of what went into it.

The main goal was to make it feel like an actual mid-90s movie, rather than a collection of unrelated AI shots. In my head this was a movie I wish Spielberg directed when I was a kid, so I tried to channel my inner 13-year-old when in doubt.

The story leans a lot into the paranoid technology movies of that period—Hackers, The Net, The Lawnmower Man, etc. Names like CyberCore, the ZX-4000 and Cyber Slayer were all meant to sound like something a screenwriter might have come up with in 1994.

Images and LoRAs

The first image I made for the project was the TV-monitor shot in the bedroom, generated in Midjourney near the end of 2024.

After that I used a little bit of everything: Flux, Z-Image, Adobe Firefly, Gemini, Grok and several others.

For the actors, I collected screenshots from their 90s movies and trained character LoRAs for each of them. I originally used Flux for most of this, but later found that Z-Image generally gave me better likenesses.

When direct generation didn’t work, I would create a lookalike, bring the image into ComfyUI and inpaint the face using the appropriate LoRA.

And sometimes I just opened Photoshop and fixed it.

I also made character sheets for the recurring characters and monsters, including both costume versions of myself. For sets like the boardroom and digitization chamber, I made reference images and rough room layouts so the shots would have some continuity.

For the creatures, I tried to think about what could realistically have been done in 1995. I often prompted for latex creatures, animatronics, puppets, miniatures or physical models—and sometimes specifically said not CGI.

I wanted them to feel more like something Stan Winston or ILM might have built than a modern digital creature.

The final color treatment was also important. I added grain, softened the image, adjusted the colors and pulled things back from the ultra-clean digital look. The footage is supposed to feel slightly faded and imperfect.

Video generation

Every generated video shot was made locally in ComfyUI or built further in After Effects.

Most of the finished trailer was generated with Wan 2.1 and Wan 2.2, although I replaced and improved several shots with MiniMax H3 during the final week.

My original plan was to film myself acting out the performances and transfer that movement onto the actors using Wan Animate.

The body movement worked surprisingly well but the faces did not.

They would gradually morph until the actors stopped looking like themselves. I tried several ways to repair them, but only one or two shots from that workflow survived.

Most of the trailer used more traditional image-to-video generations with a starting frame.

Sometimes I would generate a video mainly because I wanted the model to show me the room or character from another angle. I would grab a single frame from that result, clean it up, inpaint the face again if necessary, and then use that frame as the starting image for a completely different shot. Many times I would also grab a frame from a video which wasn’t working and then use that as an End Frame and then generate again.

Prompting video models eventually started feeling like learning another language. It took a long time to figure out how to describe blocking, timing and camera movement in a way that produced something close to what I wanted.

For MiniMax H3, I used ChatGPT and Codex to build a custom prompt builder. That let me spend less time worrying about model-specific formatting and more time thinking about the actual shot.

Green screen and compositing

A few shots are real footage of me filmed against a green screen, including:

  • Playing video games in the bedroom
  • Standing underneath the digitization lasers
  • Talking to Stallone in the desert

I filmed those in my garage or backyard. I bought costumes for both versions of the character, generated and animated the backgrounds separately, and then tried to match the lighting on myself as closely as possible.

All of the animated TV and computer monitors were composited manually in After Effects.

For those shots, I first generated a version with the screen turned off. That gave me a clean plate containing the reflections from the room on the glass.

I tracked the footage, added the animated screen underneath, and then reused parts of the original plate over the top to restore the reflections. Without that step, the screens looked like flat images pasted onto the monitors.

Voices, music and sound

All of the dialogue started with my own recorded performances.

I built custom RVC voice models for the actors, but I still performed every line myself because I wanted the timing and delivery to resemble the actual performers.

Alan Rickman has a very specific cadence, so I had to pay close attention to the rhythm of his lines.

Arnold is equally recognizable, but for different reasons.

“Down there” needed to be closer to “Down deyah.”

For the Don LaFontaine-style narrator, I went through dozens of old trailers and pulled out usable voice clips. Most needed a lot of cleanup because the narration was buried underneath music, explosions and other effects.

I also did a complete sound-effects pass. A few shots retained usable generated audio, but most of it had to be designed or sourced separately.

There was one Stallone scream during the cliff jump that I could never get the voice model to perform convincingly, so I borrowed a scream from Demolition Man.

The music was generated with Suno after weeks of attempts.

For the main action section, I found a piece of music I liked first and then edited the trailer around it.

For the final Harrison Ford moment, I wanted to hint at a classic adventure score without directly copying one. I recorded myself humming a rough melody, gave that to Suno and let it turn my bad humming into an orchestral stinger.

Making it feel like one movie

The story evolved while I was working, but I always wanted the trailer to feel like there was a complete movie behind it.

That meant thinking about continuity, character geography, setups and payoffs, when to introduce someone, when to hold back a reveal and whether one shot actually made sense next to another.

AI makes it fairly easy to generate an interesting isolated shot.

Getting dozens of shots—made with different models, months apart—to feel like they came from the same movie was the real challenge.

All told, this took around 6–8 months, mostly working on it at night after work. It was fun, but also exhausting.

It obviously isn’t perfect, and I can still see things I would change if I kept going, but eventually I had to decide it was finished.

Tech-wise I started with a 4060 Ti but decided to bite the bullet and snagged a 5090 (I also have 64GB of RAM). That helped to speed things up a ton.

I uploaded the video directly here, but there is also a YouTube version which might be higher quality

Happy to answer questions or break down any particular shot, LoRA, model, composite or workflow.

Source: Reddit

1 Like

Riportato.

3 Likes

Se l’hanno davvero fatto con Wan 2.2 è notevole.

Wan adesso sta alla 2.7 (e credo sia uscita la 3.0) e rispetto a prima è migliorato tantissimo.

1 Like

Io son onestamente molto combattuto perche’ c’e’ chiaramente un sacco di lavoro dietro comunque, non e’ che scrivi “oi fammi un trailer cusi’ cusa’ ciaone”. E’ anche uno che lavora nell’ambito che e’ esattamente quello che faccio anche io e l’AI certamente non fa il “lavoro per me” ma mi permette di arrivare a certi risultati che non avrei potuto raggiungere normalmente.

Di contro il fatto di usare attori che sono persone reali mi lascia un discreto senso di disgusto.

Il puro aspetto tecnico e’ sicuramente impressionante, ma c’e’ una parte di morale che va considerata perche’ se non la consideriamo non siamo piu’ esseri umani.

2 Likes

Allora, dal mio punto di vista trovo il risultato finale poco piacevole.

Cioè comprendo lo sforzo di piegare il generatore di immagini al proprio volere, ma mi appare evidente la ancora distanza abissale rispetto ad attori reali.

Già dopo alcuni secondi ne percepisco la resa come dire, posticcia, la metrica del parlato è ancora distante anni luce.

Insomma mi infastidisce che ne intuisco subito l’artificialità. E non è nemmeno per il fatto di vedere dei volti noti, ma proprio il fuori “fase” dello sguardo rispetto al significato delle parole proferite o la scena troppo pulita, cose così.

Secondo me è ancora una sfida tutta aperta, il verosimile è ancora riconoscibile.

Certo che il momento che non lo sarà più, cosa che richiederà secondo me un paradigma nuovo di “percezione di umano reale”, sarà divertente. Nutro ancora seri dubbi che gli attuali tools siano in grado di fare il davvero realistico. Specialmente i volti umani che dicono cose.

Cioè per farlo dovrebbero imparare tutto il ventaglio di layers empatici e come questi vengono percepiti dagli umani.

Non basta dare le sembianze di Rambo, devi essere Rambo.

Che è un paradigma totalmente differente.

In 24 fotogrammi a muzzo che prendi ci stanno una quantità di informazioni visive talmente impercettibili che ritengo davvero improbo tentare di trasferire a una AI.

Comunque il video è davvero una nerdata di quelle notevoli devo ammettere. Il che è comunque commuovente per altri motivi.

Chi di noi non ha nerdato durissimo nella propria vita.

1 Like

Si Hans, ma come ho detto io, stiamo parlando di Wan 2.2.

Ci sono dei risultati con Wan 2.7, tipo il combattimento tra Cruise e Pitt, dove tutte ste problematiche iniziano a svanire completamente.

Il resto di queste produzioni ha semplicemente un problema molto marcato: l’assenza di una regia.
Quando ad esempio ci sono quei frame di stacco con primi piani che fanno zoomout, è una tipica impostazione dell’IA che va corretta con un prompt registico marcato.

E’ praticamente il segno che le IA stanno forse levando spazio agli attori, ma stanno richiedendo l’intervento di più esperti, come sceneggiatori, registi, storyboarder, etc…

1 Like

Che e’ la cosa che sta succedendo in ogni campo btw.

La gente che panica in IT e’ gente che, mi spiace per loro, faceva solo code typing e si sentiva fiera di scrivere n-cento righe al giorno. O “vedi che soluzione elegante, risolve in 2 funzioni recursive”. Solo che fare IT engineering e’ molto piu’ di quello e se hai 20 anni di esperienza e sei a fare code typing come identita’ del tuo ruolo, e’ un problema grosso.

Il fenomeno del percepire tutto sto “slop” e’ perche’ l’utente medio di questa roba ha idee di merda e non ha minimamente nessuna preparazione o esperienza e crea roba che lui stesso reputa bella e la shara. Come in tutte le cose, quando si abbassa il livello di accesso a certe funzioni, la prima ondata e’ merda che caga merda perche’ tutti ci si buttano e si sentono fighi. Poi “passa” e si normalizza sia l’utililzzo che la qualita’.

Il problema vero con questa specifica tecnologia e’ che e’ troppo troppo troppo costosa per usi enterprise ma per uso personale sta diventando sempre piu’ efficiente. Purtroppo i vari CEO non capiscono sta cosa e sanno solo che i soldi stanno nel B2B quindi daje a creare macro-AI-warehouse, daje a fare contratti di decine di miliardi tra aziende per fornirsi a vicenda servizi.

Come sempre cazzo, il problema non e’ lo strumento, e’ l’imbecille che non sa che va usato in una maniera decente o finisci per spararti su un piede. In questo caso pero’ lo strumento spara bombe nucleari a livello sociale e ambientale che e’ un po’ un problema.

2 Likes