ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

2,461,511 results

DigitalOcean

60.2K subscribers

Latest from DigitalOcean

DigitalOcean
I made 4 LLMs fight a naval war. They signed treaties first.

Four LLMs command fleets in Flotilla, a naval resource-gathering sim. Every 100 ticks, each model rewrites the program that runs ...

9:43
I made 4 LLMs fight a naval war. They signed treaties first.

717 views

3 days ago

DigitalOcean
A fine-tuned Qwen 7B beat Llama 70B | fine-tuning vs prompting | real benchmarking 2026

I fine tuned a Qwen 2.5 7B model and put it against a Llama 3.3 70B ten times its size, same task, same 250 test items, both in ...

10:22
A fine-tuned Qwen 7B beat Llama 70B | fine-tuning vs prompting | real benchmarking 2026

269 views

3 weeks ago

DigitalOcean
Why Your vLLM p99 Latency Falls Apart in Production (and How to Fix It)

Your vLLM server benchmarked fine, then p99 latency got 5–10x worse under real traffic. The cause isn't your GPU, it's line ...

8:45
Why Your vLLM p99 Latency Falls Apart in Production (and How to Fix It)

193 views

1 month ago

DigitalOcean
We Built the Same App Twice  With and Without Kimi K3

Join the AI-Native Cloud: ...

35:32
We Built the Same App Twice With and Without Kimi K3

389 views

1 month ago

DigitalOcean
How Busy Should Your GPU Be? Picking an LLM Inference Stack

Serverless API, managed dedicated endpoint, or a GPU you run yourself — we ran the same model (Qwen3-32B on vLLM) with ...

9:08
How Busy Should Your GPU Be? Picking an LLM Inference Stack

245 views

1 month ago

DigitalOcean
When Is Serverless Inference Cheaper Than Running Your Own GPU? Real Benchmarks For 2026

Serverless inference vs running your own GPU: I benchmarked both with the same model (gpt-oss-120b) and measured the part ...

6:59
When Is Serverless Inference Cheaper Than Running Your Own GPU? Real Benchmarks For 2026

501 views

1 month ago

DigitalOcean
What is inference routing? (OpenRouter alternative)

Cut your LLM inference costs by routing each call to the cheapest model that can handle it. In this video you'll learn how inference ...

12:59
What is inference routing? (OpenRouter alternative)

656 views

2 months ago

DigitalOcean
Agent Memory Showdown

Most chatbots forget what you told them five minutes ago. So which memory architecture actually keeps an AI agent sharp at scale ...

7:03
Agent Memory Showdown

408 views

2 months ago

DigitalOcean
We Got 2x LLM Inference Speed With Three Kubernetes Settings

Scaling LLM inference on Kubernetes sounds straightforward, until you realize every pod needs 140GB of model weights, and ...

10:02
We Got 2x LLM Inference Speed With Three Kubernetes Settings

441 views

3 months ago

DigitalOcean
The Inference Economy: How Venture Is Betting on the Agentic Era

Leading investors break down the economics of scaling AI in production, from infrastructure bottlenecks to open vs. closed ...

33:09
The Inference Economy: How Venture Is Betting on the Agentic Era

287 views

3 months ago