Table of Contents
Introduction
Enhancing video footage can be a hardware-intensive process, especially when it involves AI models and tools. Certain models are only accessible through cloud-based services because their processing requirements are complex and can exceed the capabilities of even high-end desktop workstations. Fortunately, some software vendors are optimizing these complex AI models so they’re available for local processing on consumer- and professional-grade hardware.
One popular tool used in video post-production workflows for AI-generated enhancements is Topaz Video. This application is available in the cloud via its web-based interface and desktop app, offering a wide array of proprietary AI models to enhance footage. Some of their models are classified as “precise” enhancements, meaning they improve specific aspects of the footage without altering the original content (pixels) in each frame. Other models are classified as “creative” – aka generative models. Those use diffusion-based processes, such as Topaz’s Starlight models, which reconstruct each frame while maintaining temporal consistency across the video.

Each variant of the Starlight model is designed for a specific type of enhancement, and access to a given model depends on whether processing is performed locally or remotely. Those who don’t have access to a powerful workstation can opt for Topaz’s web-based platform. It offers some of the Starlight models, as well as their more complex Astra set of models, but processing is handled only through Topaz’s cloud rendering service. That incurs a cost in credits that varies by the model chosen for enhancement, and is limited in terms of what models and codecs are available for processing and export.

Topaz For Web User Video Enhancement Interface
However, those with access to a high-end workstation can use Topaz Video, which runs through NeuroServer, to process all variants of the Starlight model locally on their own hardware – without incurring the processing costs associated with Topaz for Web. Because of how new they are, we have not found any benchmark or database that shows performance differences among current-generation GPUs when processing these models locally. This prompted us to conduct our own tests with each Starlight variant, to better understand GPU performance in this space. By focusing on consumer-grade video cards from AMD and NVIDIA, this article shows which cards are compatible with each model and how they compare to each other.
Testing Methodology
Since we have not published any performance data on Starlight before, we wanted to provide details on our testing and scoring methodologies.
All testing in this article was performed manually in Topaz Video version 1.6.1, as we have yet to develop an automated script or a formal benchmark for testing the Starlight model variants. Our scope included Starlight Mini, Starlight Sharp, Starlight Fast 2, and Starlight Precise 2.5. It’s worth noting that Starlight Precise 2.6 was not available yet during our testing, so it is not included in these results, but it is now available in Topaz Video version 1.7.

Source Video Clip of Tacoma, Washington Used in Starlight Testing
The source video clip used in each test has the following specifications: 1920×1080 resolution, H.264 8-bit 4:2:0 AVC codec, wrapped in an .mp4 container at 23.976 FPS, with a runtime of 10 seconds. We chose this format because it reflects a common type of video file that users might enhance with Starlight, whether to restore archival media, repurpose highly compressed footage for a new project, or refine the pixels of already-high-quality footage.
For each Starlight variant, we processed 60 frames (approximately 2.5 seconds) from the source video clip rather than the full 10-second clip. This kept total testing time manageable, as Starlight models can take a significant amount of time to process. We also upscaled all clips by 2x during testing, since the maximum output resolution for the Starlight models is 4K (3840×2160).

Screenshot of Topaz Video UI with Starlight Model Variants
We ran five back-to-back exports to capture any variance in processing time. From there, we used Topaz Video’s log file to extract the Total Processing Time (TPT), as it’s the only metric that appears consistently across all Starlight variants. We then averaged the TPT across the five export runs and divided that time (in seconds) by 60 frames (2.5 seconds of ~24fps video) to obtain a time in seconds per frame (SPF). We chose SPF over frames per second (FPS) because it provides a more understandable metric for comparing performance across GPU models. FPS would have been measured in fractions, since even the fastest of the Starlight models takes more than a second per frame.
It’s worth mentioning that Total Processing Time reflects only the actual processing work performed on the source clip. This time does not include model loading, NeuroServer (a.k.a. NeuroStream) optimization for the system’s hardware, or exporting the final video clip. Real-world usage of these models will take somewhat longer than the results shown here. Even though our data only reflects one of the four steps involved, it is the most time-consuming – especially if you are processing longer clips.
To illustrate our process, the table below shows the presets used for testing each Starlight variant:
| Source Video Format | Frames Processed | Upscale Settings | Starlight Variant | # of Processing Runs | Seconds Per Frame (SPF) Calculation |
| 1920×1080 8bit 4:2:0 23.976 FPS | 60 Frames (2.5 seconds) | 2x | Mini Fast 2 (Local) Sharp Precise 2.5 | 5 | Average TPT of Runs 1-5 ÷ 60 frames |
For those who wish to replicate our testing, the Processing settings in Preferences were set to a max process of 1, a max memory of 100%, and a single GPU. GPU settings were optimized for single-video processing, and the default image sequence frame rate was set to match the clip, at 23.976 FPS. Because the purpose of this test was to evaluate processing performance rather than encoding performance, we used ProRes 422 LT as the encoding codec for all exports. It’s also worth noting that we did not test other codecs, input resolutions, or output settings – so processing times may differ for footage with specs different from those used in our testing. We also did not test longer clip durations, though processing time should scale directly with the number of frames queued for processing.
Screenshots of Topaz Video Settings
Our articles typically include an ‘Overall Score’ based on the geometric mean of results gathered across the tested applications or workloads, to make GPUs easier to compare. However, some of the cards we tested were not compatible with all of the Starlight variants and thus would have been excluded from this score. For that reason, we are not including an Overall Score this time. Instead, we are simply comparing the performance of current-generation graphics cards that are compatible with each model. If a GPU is not listed on the chart for a specific variant, that means it was incompatible.
Test Setup (Expandable)
Test Platform
| CPU: AMD Ryzen™ Threadripper™ PRO 9965WX |
| CPU Cooler: Asetek 836S-M1A 360mm |
| Motherboard: ASUS Pro WS WRX90E-SAGE SE BIOS Version: 1317 |
| RAM: 2x DDR5-6400 ECC Reg. 16GB (128 GB total) |
| PSU: EVGA SuperNOVA 1600W P2 |
| Storage: Samsung 980 Pro 2TB |
| OS: Windows 11 Pro 64-bit (26200) |
Software
| Topaz Video 1.6.1 |
Our test was conducted on an AMD Ryzen™ Threadripper™ PRO 9965WX-based platform. Before we chose the 9965WX, we tested a handful of CPUs to determine which would be best suited for this testing. We found that AMD’s Threadripper processor performed 10% faster than consumer-class models such as the Intel Core™ Ultra 7 270K Plus and AMD Ryzen™ 9 9950X3D2 Dual Edition. A consumer-class CPU is certainly acceptable for Topaz Video, as it’s mainly a GPU-intensive application, but we wanted to eliminate potential bottlenecks as much as possible. Those looking to get the most out of Topaz should consider a Threadripper-based workstation for both its performance and the additional PCIe bandwidth it provides.
Furthermore, we chose the workstation-class Threadripper PRO over the standard Threadripper because the motherboard’s additional PCIe bandwidth allows us to test performance scaling from one to four GPUs. We will be looking at that in an upcoming article.
To better understand how current-gen consumer GPUs perform when running different Starlight models locally, we tested several models from both AMD and NVIDIA. Intel’s Arc B580 was excluded, as it is not compatible with any of these models. We used the latest available drivers at the time testing began, but newer driver versions and software updates have since been released – so the results reflected here may no longer be fully current.
Starlight Consumer GPU Performance Analysis
Each variant of the Starlight model is designed for a specific purpose. While all of them perform some level of enhancement, certain variants are optimized for processing speed, while others place a heavier emphasis on detail recovery and enhancement to produce higher-quality results. Depending on the user’s needs, some workflows may use a single Starlight variant, while others may use a combination to achieve the best results.
With that in mind, the charts below show which of the tested GPUs are compatible with each Starlight variant and how they perform compared to each other.
Starlight Mini

Our first analysis is with Starlight Mini, the first variant that Topaz Video made available for local processing. This model is designed to restore heavily degraded footage, such as film and other archival footage.
Starlight Mini only works with NVIDIA graphics cards. Among the cards we tested, the GeForce RTX 5090 is the top performer, processing a frame in about 8 seconds — roughly 60% faster than the GeForce RTX 5080, which took about 13 seconds per frame. The GeForce RTX 5070 Ti trailed the 5080 by only about 10% in Starlight Mini processing time, a gap small enough that price may be a substantial factor in choosing between them. Further down the chart, the GeForce RTX 5070 and GeForce RTX 5060 Ti see much longer processing times: the RTX 5070 took about 20 seconds to process a frame, while the RTX 5060 Ti was another 26% slower at nearly 27 seconds.
Those using Starlight Mini to restore archival material without strict time constraints can get by with a slower, more budget-friendly GPU. That said, a faster card is worth considering for anyone working with longer clips or larger volumes of footage, where those extra seconds per frame will add up over time.
Starlight Sharp

Moving on to Starlight Sharp, this variant is designed to enhance low-resolution footage, with a specialty in recovering fine facial features and details in small, distant faces within the frame. As with Starlight Mini, NVIDIA graphics cards were the only ones tested, as AMD and Intel cards are incompatible with this variant. Starlight Sharp requires at least 16 GB of VRAM to function, so the GeForce RTX 5070 was also excluded from this test.
The RTX 5090 is once again the top performer, taking just over a second to process a single frame. The RTX 5080 and 5070 Ti are next and are nearly identical in performance, processing a frame in under 3 seconds. The 5060 Ti was the slowest performer, taking about 5 seconds to process a single frame.
If Starlight Sharp is the main variant a user plans to use, the RTX 5070 Ti offers the best cost-to-performance ratio, though those needing strong performance across multiple Starlight variants or dealing with a lot of footage may still want to aim for the RTX 5090.
Starlight Fast 2 (Local)

The next Starlight variant we tested is Starlight Fast 2, which is designed to work with higher-quality footage and is optimized for better processing speeds on consumer-grade NVIDIA GPUs.
Like Starlight Mini and Sharp, AMD and Intel GPUs are not compatible with this model variant. Starlight Fast 2 also requires at least 16 GB of VRAM, so the GeForce RTX 5070 was again excluded from this test. When looking at the chart above, we see some interesting results: the RTX 5090 is the top performer, taking just over one second to process a single frame, while the RTX 5060 Ti, 5070 Ti, and 5080 all take between 14 and 16 seconds to process a single frame.
Since this model is intended to be a “fast” iteration of Starlight, we suspect performance may be bottlenecked when VRAM is insufficient. In the results above, the RTX 5060 Ti actually edges out the RTX 5080, despite the 5080 being the stronger card on paper. This result could point to a memory bottleneck, where a faster GPU like the 5080 is more affected by VRAM constraints relative to its compute potential, while a comparatively slower card like the 5060 Ti is less affected by the same limitation. However, given that VRAM capacities on most current-generation hardware are capped at 16 GB, we aren’t able to test whether capacities above that threshold would mitigate this performance bottleneck.
Whatever the cause, the GeForce RTX 5090 is the clear choice for those seeking the best possible performance with Starlight Fast 2. It is more than ten times faster than the other cards, far eclipsing its higher price. For those wanting even more VRAM, we will be looking at performance in NVIDIA’s professional-grade cards soon.
Starlight Precise 2.5

The final variant we tested is Starlight Precise 2.5, which works best with high-quality video and is not well-suited for heavily degraded footage. It is designed to improve AI generated content that may have a soft, plastic, or slightly artificial look. Additionally, it is the only model in this group that supports AMD graphics cards.
Our results show that the RTX 5090 is once again the top performer, taking 5 seconds to process a single frame. Second-best is the RTX 5080, which was 37% slower, followed by the RTX 5070 Ti. The ~20% gap between those cards may be small enough to make the less expensive model more appealing for many users. The remaining NVIDIA GPUs were substantially slower, with the RTX 5070 taking nearly 14 seconds to process a single frame, while the RTX 5060 Ti is 20% slower than the 5070.
As mentioned, Starlight Precise 2.5 is also compatible with the AMD Radeon RX 9070 XT, the only non-NVIDIA GPU we tested. However, its performance was underwhelming. The 9070 XT took just over 54 seconds to process a single frame — more than three times longer than the slowest NVIDIA card in this comparison. As such, we wouldn’t recommend an AMD GPU for Starlight; NVIDIA GPUs deliver much better performance and are compatible with all model variants (assuming VRAM requirements are met).
Which GPU Is Best for Topaz Video’s Starlight?
Topaz Video’s Starlight models can produce great results, but they’re processing-intensive and, depending on the variant used, can take multiple seconds to process a single frame. That time may be further amplified if a workflow combines different models to achieve the best results. Older or slower hardware is fine for occasional use, but for time-sensitive projects or workflows that rely on multiple Starlight variants, even a modest hardware upgrade can meaningfully reduce processing time.
NVIDIA’s GeForce RTX 5090 is the clear choice for anyone limited to a single consumer-class GPU who wants to maximize performance across all Starlight variants. However, the 5090 is quite expensive in today’s market, and thus may not fit into every budget. Determining which graphics card is best for local processing with Starlight ultimately depends on factors such as which variants are used, the number of frames to be processed, and how much time can reasonably be spent waiting for Topaz Video to process them.
For those who want the same flexibility across all Starlight variants, but can’t justify the 5090’s price, the pool of compatible GPUs narrows to the RTX 5080, 5070 Ti, and 5060 Ti (16GB version). Those looking to maximize performance within that group should consider the 5080, though the 5070 Ti is often very close in performance for the more budget-constrained.
The RTX 5060 Ti is the most affordable of the three, but its processing times are extremely long. We’d only recommend it for those who aren’t time-constrained and are working with shorter-duration video clips. Moreover, the 16GB version of that card is also very hard to find now — so it may not be an option much longer. The 8GB version wouldn’t be sufficient for most of these models, making it a non-starter.
The RTX 5070’s limited 12GB of VRAM means it is only compatible with Starlight Mini and Precise 2.5, making it a poor fit for those looking to use all of the Starlight variants. Lastly, while the AMD Radeon RX 9070 XT is technically compatible with Starlight Precise 2.5, its performance was low enough that we don’t think it is worth considering. Topaz Video users are better off choosing an NVIDIA GPU with 16GB or more if they want access to all Starlight variants and Topaz’s proprietary classical GAN-based models.
Many users will want to run Starlight alongside Topaz’s classical models, so this article can be used in tandem with our previous posts that used Topaz Video’s internal benchmark. We published one covering consumer-grade cards and another looking professional-grade cards. Up next, we’ll test professional-grade GPUs with the Starlight variants, using the same testing methodology as this article, followed by multi-GPU testing with a smaller set of Topaz’s proprietary models across both consumer- and professional-grade cards.

