ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

2,116,905 results

DigitalOcean

60.2K subscribers

Latest from DigitalOcean

DigitalOcean
Why Your vLLM p99 Latency Falls Apart in Production (and How to Fix It)

Your vLLM server benchmarked fine, then p99 latency got 5–10x worse under real traffic. The cause isn't your GPU, it's line ...

8:45
Why Your vLLM p99 Latency Falls Apart in Production (and How to Fix It)

68 views

3 days ago

DigitalOcean
We Built the Same App Twice  With and Without Kimi K3

Join the AI-Native Cloud: ...

35:32
We Built the Same App Twice With and Without Kimi K3

215 views

5 days ago

DigitalOcean
How Busy Should Your GPU Be? Picking an LLM Inference Stack

Serverless API, managed dedicated endpoint, or a GPU you run yourself — we ran the same model (Qwen3-32B on vLLM) with ...

9:08
How Busy Should Your GPU Be? Picking an LLM Inference Stack

207 views

10 days ago

DigitalOcean
When Is Serverless Inference Cheaper Than Running Your Own GPU? Real Benchmarks For 2026

Serverless inference vs running your own GPU: I benchmarked both with the same model (gpt-oss-120b) and measured the part ...

6:59
When Is Serverless Inference Cheaper Than Running Your Own GPU? Real Benchmarks For 2026

440 views

3 weeks ago

DigitalOcean
What is inference routing? (OpenRouter alternative)

Cut your LLM inference costs by routing each call to the cheapest model that can handle it. In this video you'll learn how inference ...

12:59
What is inference routing? (OpenRouter alternative)

565 views

1 month ago

DigitalOcean
Agent Memory Showdown

Most chatbots forget what you told them five minutes ago. So which memory architecture actually keeps an AI agent sharp at scale ...

7:03
Agent Memory Showdown

387 views

1 month ago

DigitalOcean
We Got 2x LLM Inference Speed With Three Kubernetes Settings

Scaling LLM inference on Kubernetes sounds straightforward, until you realize every pod needs 140GB of model weights, and ...

10:02
We Got 2x LLM Inference Speed With Three Kubernetes Settings

399 views

2 months ago

DigitalOcean
The Inference Economy: How Venture Is Betting on the Agentic Era

Leading investors break down the economics of scaling AI in production, from infrastructure bottlenecks to open vs. closed ...

33:09
The Inference Economy: How Venture Is Betting on the Agentic Era

259 views

2 months ago

DigitalOcean
Your Model Doesn't Matter. Your Infrastructure Does.

Everyone has access to the same models. So what actually matters? It's everything around them – routing requests to the right ...

20:53
Your Model Doesn't Matter. Your Infrastructure Does.

327 views

2 months ago

DigitalOcean
Open by Design: How NVIDIA and DigitalOcean Are Building the Stack for the Always-On Agentic Era

Kari Briski, VP Gen AI, NVIDIA, and Salman Paracha, SVP AI, DigitalOcean discuss why AI-native teams are demanding ...

30:38
Open by Design: How NVIDIA and DigitalOcean Are Building the Stack for the Always-On Agentic Era

226 views

2 months ago

Guia Español
Como Usar DigitalOcean | Como Funciona DigitalOcean (2026)

Como Usar DigitalOcean | Como Funciona DigitalOcean (2026) En este vídeo, también hemos cubierto Aprende cómo usar ...

8:03
Como Usar DigitalOcean | Como Funciona DigitalOcean (2026)

1,592 views

5 months ago