Instagram story viewer> @rick.theengineer> Posts
407.7K
followers
238
following
Tech explained by the smartest man in the universe πŸ§ͺ(Unofficial Rick parody/fan account)
VIDEOS MADE WITH brainrotshorts.com
πŸ‘‡ The AI I Use πŸ‘‡
POSTS STORIES REELS TAGGED
Download All
One physical server can run 5, 10, even 50 "computers" at the same time. Here's how πŸ‘‡

For years, companies ran one app per server. Most of the day that machine sat at around 10% of its power, and you still paid for 100% of the hardware, the space and the electricity.

Virtualization fixes that with one piece of software: the hypervisor.

Think of the server as an apartment building:
🏒 The building = the real hardware (CPU, memory, storage, network)
πŸ”‘ The landlord = the hypervisor
🏠 The apartments = virtual machines (VMs)

The hypervisor cuts the real hardware into slices. Every VM gets its own CPU, memory, disk and operating system, and honestly believes it's a complete computer. One runs Linux for your website, another runs Windows, and none of them know the others exist.

A few things that surprise people:
β†’ VMs take turns on the real CPU cores, switching so fast that each one feels alone on the machine
β†’ One VM can't read another VM's memory. There are walls in between
β†’ A VM's whole hard drive is often just one big file on the real machine
β†’ Modern Intel and AMD chips have virtualization built in, so VMs run at almost full speed

That's the cloud. When you rent a server from AWS, you usually get a VM, not a whole machine. Your server is living in a building full of strangers.

And no, VMs are not containers. A VM carries its own complete operating system (its own kitchen). Containers share the host's kernel (one shared kitchen). Containers are lighter and start in seconds. VMs are heavier, but the walls are thicker.

One machine underneath. Many machines on top.

πŸ’Ύ Save this for your next system design interview, and send it to the friend who still thinks "the cloud" is magic.

Did you know your cloud server was probably a VM? Tell me below πŸ‘‡

#virtualization #cloudcomputing #softwareengineering #devops #techexplained by @rick.theengineer
24
16 hours ago
Download
Most teams don't break production because they're bad at code. They break it because a human had to remember every step.

SSH into the server. Pull the latest code. Install. Build. Copy files. Restart. Hope nothing breaks.
Forget one step, like the database migration, and the site goes down with a 502 error.

That's the problem CI/CD solves. It has two halves:

CI (Continuous Integration) answers one question: is this code safe to merge?
Every push runs the same pipeline automatically. It installs dependencies, runs the linter, runs unit and integration tests, builds the app and scans for vulnerabilities. If one check fails, the pipeline turns red and that code can't be merged until it's fixed.

CD answers a different question: can we safely release it?
The tested code becomes a release artifact (often a Docker image). It goes to staging first, where more checks run, and then on to production.

Here's the part most people mix up:
β†’ Continuous DELIVERY: a human still clicks "deploy" before production.
β†’ Continuous DEPLOYMENT: if every check passes, it ships automatically. No human needed.

And when something still goes wrong, good pipelines don't just ask "did the server start?" They ask "are health checks passing?" They roll out to 5% of users first (a canary), then 25%, 50%, 100%, and they roll back the moment error rates spike.

Everything in this video comes from a real run: real commits, real test results, a real failing test (expected $3.56, got $4), and a real GitHub Actions workflow file.

The real point of CI/CD isn't speed. It's that every release follows the exact same steps, every single time.

Watch it to the end: the bug that brings production down at the end is the step that got forgotten at the start πŸ‘€

Save this for your next system design interview, and send it to the friend who still deploys over SSH.

Which side does your team run: delivery or deployment? πŸ‘‡

#cicd #devops #githubactions #softwareengineering #coding by @rick.theengineer
111
2 days ago
Download
For 60 years, we threw away the most expensive part of every rocket. Every single time. πŸš€

Imagine flying New York to London, landing, and then pushing the plane into the ocean. That's how spaceflight worked until December 21, 2015, when a Falcon 9 booster flew back from space and landed standing up for the first time.

Here's what actually happens after liftoff:

1️⃣ About 2.5 minutes in, the booster separates at around 6,000 km/h. The second stage keeps going to orbit.
2️⃣ Small nitrogen thrusters flip it around in near-vacuum.
3️⃣ The boostback burn turns it toward the landing zone.
4️⃣ The entry burn slows it down so it survives the air getting thick again.
5️⃣ Titanium grid fins (the waffle-looking things) steer it through the atmosphere.
6️⃣ One engine fires a final time, the legs open, and it lands. Nobody is flying it. Computers re-plan the descent in a fraction of a second.

Why bother? The booster is roughly 60% of the rocket's cost. Fly it again and that cost gets spread across many launches. One booster has now flown 37 times, and the fastest turnaround has been 9 days.

But reusable isn't free. Coming back takes fuel that could have carried payload: about 17.5 tonnes to low Earth orbit when the booster lands, versus 22.8 tonnes when it's thrown away. Every mission is a trade-off.

Reusable rockets don't make spaceflight simple. They make it more like aviation: launch, land, refuel, fly again. ✈️

Which part surprised you most, the flip, the fins, or the landing? Tell me below πŸ‘‡ and save this for the next launch you watch.

#reusablerockets #spacex #falcon9 #rocketscience #spaceexplained by @rick.theengineer
47
3 days ago
Download
Your password is a secret you've already handed to every website you use. A passkey never leaves your phone. πŸ”

Here's what actually happens behind the login screen πŸ‘‡

With a password, you and the website share the same secret. You type it, it travels to the server, and the server checks it against a stored copy (a hash). That's why passwords get reused, guessed, leaked in data breaches, and typed into fake sites.

A passkey works completely differently:
β†’ Your phone creates a key pair: one private, one public
β†’ The website only ever stores the PUBLIC key
β†’ When you log in, the website sends a random challenge
β†’ Your phone signs it with the PRIVATE key, unlocked by Face ID, Touch ID or your PIN
β†’ The website checks the signature with the public key, and you're in

The private key is never sent anywhere. Your face never leaves your phone either.

And the best part? Phishing stops working. A fake "paypaI.com" (with a capital I) looks identical to the real one, but your passkey is cryptographically tied to the real domain. On the fake site, your phone has nothing to offer. Nothing to type, nothing to steal.

If a website gets breached? Attackers get a public key that can't log anyone in.

The catch: if you lose every device, account recovery matters. That's why iCloud Keychain and Google Password Manager sync your passkeys (end-to-end encrypted), and why we'll live in a mixed world of passwords, passkeys and 2FA for a while.

The easiest way to remember it:
A password says "I know the secret."
A passkey says "I can prove it, without ever revealing it."

Every key, signature and hash in this video is real, generated and verified for this explainer. Pause and check πŸ‘€

πŸ’¬ Have you switched any accounts to passkeys yet? Tell me which ones below.
πŸ“Œ Save this for the next time someone asks you what a passkey is.

#passkeys #cybersecurity #passwords #techexplained #infosec by @rick.theengineer
39
4 days ago
Download
Your photo takes one second to reach Tokyo. This is everything that happens in that second πŸ‘‡

Lena taps send on a picture of her cat. Before her brother sees it, the photo goes DOWN seven layers of her phone, ACROSS the planet, and UP seven layers of his.

7 β€” Application: the app asks to send the photo (HTTP, DNS)
6 β€” Presentation: it gets compressed (7.68 MB β†’ 313 KB) and encrypted
5 β€” Session: her phone and the server open a conversation
4 β€” Transport: TCP cuts it into 215 numbered pieces and re-sends any that get lost
3 β€” Network: each piece gets the destination IP address
2 β€” Data link: plus a MAC address to reach the home router first
1 β€” Physical: it becomes radio waves, electrical pulses, then light in fibre under the ocean

Routers read the IP label hop by hop all the way to Tokyo, where every layer runs in reverse and his app shows a cat.

Every layer wraps the data in its own label, a package inside a package inside a package. That's encapsulation.

And when a photo never arrives, the layers tell you where to look:
no signal β†’ layer 1
can't reach the router β†’ layer 2
lost on the way β†’ layer 3
pieces missing β†’ layer 4
app error β†’ higher up

The real internet runs on the simpler TCP/IP model, but OSI is still the best map of the journey.

Save this for your next networking exam or interview πŸ“Œ
Which layer surprised you the most? Tell me in the comments πŸ‘‡

#osimodel #networking #computerscience #cybersecurity #learntocode by @rick.theengineer
348
6 days ago
Download
Every app on your phone runs on a database. But not the same one. πŸ“±

Your bank, your shopping app, your login, your ride app, your search bar, your fitness watch, your music app. Seven apps, seven completely different kinds of data.

Here's the cheat sheet πŸ‘‡

🏦 Bank β†’ Relational (Postgres, MySQL)
Linked tables and all-or-nothing transfers. We actually killed a transfer halfway through. The balance didn't move by a single cent.

πŸ›οΈ Shop β†’ Document (MongoDB)
A T-shirt has a size and a color. A laptop has a battery and a processor. Each product keeps only the details it needs.

πŸ”‘ Login β†’ Key-Value (Redis)
Like a coat check. Hand over a ticket, get your coat back. We measured it at 0.015 ms per lookup.

πŸš— Rides β†’ Wide-Column (Cassandra)
Millions of location pings that never stop. Spread across many machines, and it keeps going even when one dies. The catch: you plan your questions in advance.

πŸ” Search β†’ Search Engine (Elasticsearch)
Type "blak watrproof runing shoes" and still get the right result. It works like the index at the back of a book.

⌚ Fitness β†’ Time-Series (Timescale, InfluxDB)
One heart-rate reading per second is 604,800 readings a week. Every value is tied to a time.

🎡 Music β†’ Vector (Pinecone, Weaviate, pgvector)
Why does the next song just feel right? It's found by meaning, not by matching words. The same idea powers AI chatbots.

The twist: most apps don't need all of these. Postgres alone can store documents, search text and even store vectors.

So start simple. Postgres for your core data, Redis when you need speed, and the rest only when a real problem shows up.

The best database isn't the most powerful one. It's the one that fits the problem you actually have.

πŸ’Ύ Save this for your next project
πŸ’¬ Which one does your app use? Drop it below πŸ‘‡

#database #systemdesign #softwareengineering #programming #backend by @rick.theengineer
28
6 days ago
Download
The AI ranked #1 might still be the wrong model for your work. πŸ‘€

A benchmark score answers one question: β€œHow well did this model perform on this particular test?”

It doesn’t answer: β€œIs this the best AI for everything?”

Different tests measure different abilities:
β†’ AIME tests math reasoning.
β†’ GPQA tests advanced science knowledge.
β†’ SWE-bench tests whether a model can resolve real software issues.
β†’ Agentic benchmarks test whether it can use tools to complete a task.

Then come the details that leaderboard screenshots leave out.

Did the model encounter similar questions during training? Does a tiny score difference actually matter? How much time, compute and money did that result require?

Consider this hypothetical comparison:

Model A: 92% accuracy, 30 seconds, $1 per task.
Model B: 89% accuracy, 2 seconds, $0.05 per task.

Which one wins?

That depends on what you’re buildingβ€”and what a mistake costs you.

Before choosing your next model, test it on your own code, data and workflows.

Save this for the next launch claiming β€œbest AI yet.”

What matters most to you: accuracy, speed or cost?

#aibenchmarks #artificialintelligence #aitools #machinelearning by @rick.theengineer
17
7 days ago
Download
Every answer from ChatGPT or Claude is a fight over 80 GB. 🧠⚑

Training a model happens once. Serving it to millions of people, every second of every day, is the part nobody sees. Here's what actually happens after you hit send:

1️⃣ The model doesn't fit.
A 100-billion-parameter model is about 200 GB at 16-bit. One NVIDIA H100 holds 80 GB. So the model gets split across several GPUs, and they have to talk over NVLink for every single token.

2️⃣ Your prompt is read all at once (prefill). The answer is not.
Every word you see is predicted one token at a time. Token 100 literally cannot exist before token 99.

3️⃣ One GPU, dozens of people.
Inference servers like vLLM batch users together, and every step writes one token for everyone in the batch. When someone finishes, a new user takes their lane on the very next step. That scheduling alone can double what the same GPU gets done.

4️⃣ The hidden memory hog: the KV cache.
The model keeps notes on every token of your conversation. For Llama 3.1 70B, one full 128K-token conversation needs about 43 GB of cache. That's over half a GPU for one chat.

5️⃣ The tricks that make it affordable:
β†’ quantization (16 bits down to 8 or 4)
β†’ tensor + pipeline parallelism
β†’ mixture of experts (Mixtral holds 46.7B parameters but uses 12.9B per token)

6️⃣ Why words stream in.
Tokens are sent the moment they're generated, so a 20-second answer feels instant. What matters: time to first token, and how fast the rest follow.

Training builds the brain. Inference is the infrastructure problem of letting millions of people use it at once without spending a fortune every second.

πŸ’¬ Which part surprised you most? Drop a number 1–6 πŸ‘‡
πŸ“Œ Save this for the next time someone says AI is "just an API call."

#aiinference #llm #machinelearning #nvidia #softwareengineering by @rick.theengineer
26
9 days ago
Download
Nobody ever wrote the code that lets ChatGPT write Python. Not one line of it. 🀯

Here's how a large language model actually gets built, step by step:

1️⃣ Data. Trillions of words of books, web pages, code and forums, then cleaned: duplicates removed, spam filtered, broken text stripped. One open dataset, FineWeb, is 15 trillion tokens of web text.

2️⃣ Tokens. The model never sees words. "Unbelievable" becomes 4 pieces: Un Β· bel Β· iev Β· able. Each piece becomes a number, and each number becomes a vector of hundreds of numbers.

3️⃣ The transformer. Stacks of layers, and inside each one, attention. In "the programmer fixed the server because it crashed", a real GPT-2 attention head sends 54% of "it"'s attention straight to "server".

4️⃣ Pre-training. One task, repeated trillions of times: predict the next token. Guess wrong, measure how wrong, nudge the weights a tiny bit downhill. Meta trained Llama 3 on up to 16,000 GPUs.

5️⃣ Post-training. A pre-trained model is just a very powerful autocomplete. Fine-tuning on good examples, human feedback and checkable rewards (did the math match? did the code pass?) turns it into an assistant.

6️⃣ Serving. A 405-billion-parameter model needs about 810 GB just for its weights. One GPU holds 80. So it gets split, compressed and batched so millions of people can use it at once.

Every example in the video comes from a real model: the tokens, the attention, even the loss landscape.

The weird part? Writing code, explaining physics, translating French: none of it was programmed. It all emerges from one tiny objective, repeated at enormous scale: predict what comes next, and get slightly less wrong every time.

Which step surprised you most? Drop the number πŸ‘‡

Save this for the next time someone calls AI "just autocomplete."

#llm #artificialintelligence #machinelearning #chatgpt #techexplained by @rick.theengineer
282
10 days ago
Download
Load balancer. Cache. Queue. CDN. Replica. Shard. 🧠

If those sound like a pile of random buzzwords, this one's for you.

Here's the secret: nobody sits down and designs a system with 9 components. You start with ONE server, and every new piece gets added because something broke.

Here's the whole thing in order πŸ‘‡

1️⃣ One server can't keep up β†’ add more servers + a load balancer
2️⃣ Every server hammers the same database β†’ add a cache (Redis)
3️⃣ Reads keep growing β†’ add read replicas
4️⃣ One database can't hold it all β†’ shard the data
5️⃣ Photos and videos bloat the database β†’ move them to object storage (S3 / R2)
6️⃣ Users far away wait forever β†’ put a CDN in front
7️⃣ Uploads make users wait on slow jobs β†’ add a queue + background workers
8️⃣ Servers crash, regions disappear β†’ build in redundancy
9️⃣ You can't see what's breaking β†’ add observability (logs, metrics, traces)

And the biggest beginner mistake? Adding Kafka, Kubernetes, microservices and 10 databases before you have a real problem. Every component has a cost:
β€’ caches go stale
β€’ queues deliver twice
β€’ replicas lag
β€’ shards make queries harder
β€’ microservices fail over the network

So the real skill isn't memorizing tools. It's asking one question:

πŸ‘‰ "What problem forced us to add this?"

Start simple. Find the bottleneck. Add the smallest fix. Repeat.

πŸ’Ύ Save this for your next system design interview

πŸ’¬ Which of the 9 steps did you learn the hard way? Tell me in the comments πŸ‘‡

#systemdesign #softwareengineering #backenddevelopment #codinginterview #learntocode by @rick.theengineer
260
11 days ago
Download
Grandpa & Grandson πŸ§ͺπŸ›Έ

Our little tribute to the iconic Hotel Lobby.
Shoutout to Quavo & Takeoff πŸ•ŠοΈ for the original. by @rick.theengineer
27
12 days ago
Download
your app is asking the server "anything new yet?" every second… and 99% of the time the answer is no πŸ™ƒ

there are 3 better ways to get fresh data, and they solve 3 completely different problems:

⚑ websockets: a phone call
one connection that stays open, and both sides can talk anytime.
β†’ chat apps, multiplayer games, live cursors, trading dashboards

πŸ”” webhooks: "call me when the package arrives"
no connection at all. when something happens, one server sends a normal http request to another.
β†’ stripe payments, github pull requests, shopify orders

πŸ“» server-sent events (sse): live radio
the browser tunes in once, and the server keeps streaming updates one way.
β†’ chatgpt typing its answer word by word, live feeds, progress bars

the trade-offs nobody mentions:
β€’ websockets are powerful, but thousands of open connections get expensive to run
β€’ webhooks can fail, so you need retries, signature checks and idempotency
β€’ sse is simpler to run, but it's one-way. if the client needs to talk back, use websockets

easy way to remember it:
both sides talking β†’ websockets
server tells another server β†’ webhooks
server streams to the browser β†’ sse

watch it twice, the second time you'll catch the details in the devtools panels πŸ‘€

πŸ’Ύ save this for your next system design interview
πŸ’¬ which one does your app actually use? tell me below πŸ‘‡

#webdevelopment #systemdesign #programming #softwareengineering #websockets by @rick.theengineer
28
12 days ago
Download
×

Download all media on this page

Photos Videos
back to up