WEBVTT

00:00.000 --> 00:12.000
The next talk is a really interesting one.

00:12.000 --> 00:21.000
I saw a long version of it at Aussie Stepsko where I presented this and now the

00:21.000 --> 00:25.000
evolved shortened version is here for all of us.

00:25.000 --> 00:27.000
I think what you're doing is really cool.

00:27.000 --> 00:31.000
Like sending people your talk.

00:31.000 --> 00:35.000
Maybe I can send them the first version now.

00:35.000 --> 00:40.000
I am from Putin is going to present about Luka Mahno's secrets.

00:40.000 --> 00:45.000
And then I just wrote keywords because we are bootstraping trust and then a bunch of letters.

00:45.000 --> 00:48.000
Right, UK, ITPM and spiffy.

00:48.000 --> 00:52.000
Also all the abbreviations all the stuff.

00:52.000 --> 00:55.000
But like really combined in a thoughtful sensible

00:55.000 --> 00:58.000
method to do something interesting that's really hard to do.

00:58.000 --> 01:00.000
I have fun.

01:00.000 --> 01:03.000
Thank you very much for presenting a round of applause for Ariana please.

01:03.000 --> 01:05.000
Thank you.

01:09.000 --> 01:17.000
Yeah, so as many of you do, I have a home lab at home with like a media server and a web server for my website.

01:17.000 --> 01:25.000
And yeah, one day I wanted to add monitoring to my home lab.

01:25.000 --> 01:27.000
So get like page.

01:27.000 --> 01:32.000
Yeah, okay, get page when my website is down or my media server is down.

01:32.000 --> 01:38.000
And that ended up in an enormous rabbit hole where I try to invent my own clouds.

01:38.000 --> 01:41.000
And this is what I'm going to talk about.

01:42.000 --> 01:47.000
So yeah, Nick, so as a super cool, I think a lot of you already use Nick, so as so kind of not explain what it is.

01:47.000 --> 01:49.000
But yeah, in a nutshell, this is my home lab.

01:49.000 --> 01:54.000
I have a web server, mini-dl in a server for me, fias and alert manager.

01:54.000 --> 01:57.000
And they run on different nodes.

01:57.000 --> 02:01.000
And they all, the defaults run on all of the nodes.

02:01.000 --> 02:05.000
So all the nodes run up for me, fias, no to export there.

02:05.000 --> 02:12.000
And then I set up a alert manager config to scrape all my nodes.

02:12.000 --> 02:16.000
And send me an alerts to my phone if any of the servers are down, right?

02:16.000 --> 02:24.000
And then it sends a nice little, it sends a thing to this web hook that slides, which is hooked up to my phone.

02:24.000 --> 02:28.000
And then I get a page when my media server is down.

02:28.000 --> 02:31.000
Okay, this all, all good, all nice.

02:31.000 --> 02:35.000
The problem is, all my stuff is open source and this link is on my GitHub.

02:35.000 --> 02:41.000
So people started paging my phone with funny messages.

02:41.000 --> 02:48.000
Anyone in my house can talk to alert manager and send alerts as well.

02:48.000 --> 02:52.000
So when I had somebody over who realized that they started sending funny things to my phone.

02:52.000 --> 02:58.000
Also, anybody can scrape my metrics for my servers, which also don't want.

02:58.000 --> 03:01.000
So I thought, well, I need another component in my home lab.

03:01.000 --> 03:03.000
This will solve all my problems.

03:03.000 --> 03:07.000
So I deployed OpenBow, which used to be as a fork of vault.

03:07.000 --> 03:10.000
And I thought, what if I store the web hook secret in OpenBow?

03:10.000 --> 03:13.000
Then nobody can access it, right?

03:13.000 --> 03:16.000
So, yeah, set up an OpenBow server.

03:16.000 --> 03:18.000
Awesome.

03:18.000 --> 03:24.000
So yeah, we just configure now alert manager to fetch the secrets from OpenBow.

03:24.000 --> 03:26.000
And then it's not in my repo anymore.

03:26.000 --> 03:28.000
But then I thought, okay, but how do I talk to OpenBow?

03:28.000 --> 03:31.000
I need a secret for that as well, right?

03:31.000 --> 03:37.000
So, turns out, yeah, it's turtles all the way down.

03:37.000 --> 03:42.000
I just moved to goalpost and I still have the same problem.

03:42.000 --> 03:51.000
Yeah, how do I have to make trust between my alert manager and my OpenBow so that I can get my web hook secret?

03:51.000 --> 03:59.000
So if you run on GCP or AWS, your turtle is usually cloud-shaped like this, right?

03:59.000 --> 04:01.000
And they solved all these things for you.

04:01.000 --> 04:03.000
You have like GCP service accounts.

04:03.000 --> 04:04.000
So you have a service account.

04:04.000 --> 04:07.000
You're a Kubernetes cluster or you have an IAM role.

04:07.000 --> 04:14.000
Or if I get a ID token, you just write the IAM policy and you can access secrets on your secret manager.

04:14.000 --> 04:16.000
And there's no secrets for you to manage, right?

04:16.000 --> 04:18.000
They take care of all of that.

04:18.000 --> 04:24.000
But we don't have that in my home lab.

04:24.000 --> 04:28.000
So it's like, mom, can we have cloud workload identity at home?

04:28.000 --> 04:37.000
And I started Googling and then I found this project called Spiffy, which stands for secure production identity framework for everyone.

04:37.000 --> 04:46.000
Which I think is started by some Googlers who basically wanted to take the stuff that they have internally at Google and make an open source product out of it.

04:47.000 --> 04:54.000
And it's a set of standards to talk about workload identity in a economic way.

04:54.000 --> 04:57.000
And it does this through X509 certificates.

04:57.000 --> 05:00.000
Like the ubiquitous, also for supports them.

05:00.000 --> 05:08.000
If we find a way that we can give short lift X509 certificates to workloads, then workloads can talk to each other.

05:08.000 --> 05:11.000
We need to get that X509 certificate from somewhere.

05:11.000 --> 05:14.000
So we're going to move to goal posts once more.

05:14.000 --> 05:16.000
Let's go get through there.

05:16.000 --> 05:21.000
Yeah, this is what's such an X59 certificates for my alert manager looks like.

05:21.000 --> 05:24.000
It was issued by some CA.

05:24.000 --> 05:28.000
It's very short lift and it's only valid for alert manager.

05:28.000 --> 05:30.000
And now the question is, okay, great.

05:30.000 --> 05:35.000
We use X59 certificates with where do I get that thing from, right?

05:35.000 --> 05:40.000
So there is a reference implementation called Spire, which is a certificate authority.

05:40.000 --> 05:43.000
That you can sell at host.

05:43.000 --> 05:50.000
And they implement like support for Azure, AWS, GCP and integrate with their service account stuff.

05:50.000 --> 05:56.000
But the cool thing is they also integrate with trusted platform modules, TPMs.

05:56.000 --> 06:02.000
So you can bootstrap your CA by delegating trust through your TPM.

06:02.000 --> 06:05.000
And all my devices at home have TPMs.

06:05.000 --> 06:07.000
So that kind of solves my bottom turtle.

06:07.000 --> 06:10.000
There's some unique secret in each TPM in my device.

06:10.000 --> 06:16.000
I can use that to identify each device and get a certificate from the CA,

06:16.000 --> 06:21.000
and then they can talk to each other.

06:21.000 --> 06:27.000
So the plan was, I will install Spire on my home land,

06:27.000 --> 06:30.000
run the Spire agent on each of the nodes.

06:30.000 --> 06:33.000
And then I will configure OpenBow for me,

06:33.000 --> 06:38.000
and alert manager to retrieve X59 certificates from Spire.

06:38.000 --> 06:43.000
And then we configure all these services to use TLS to talk to each other.

06:43.000 --> 06:45.000
And then nobody should be able to snoop.

06:45.000 --> 06:48.000
I should be able to see where it's from OpenBow.

06:48.000 --> 06:53.000
And I can get paged without getting funny messages from people.

06:53.000 --> 06:59.000
So yeah, how do you install Spire on your NYXOS home land?

06:59.000 --> 07:02.000
There's a pool request to open now in NYXPACHS.

07:02.000 --> 07:04.000
There's a Spire module now.

07:04.000 --> 07:11.000
And in this case, I set up Spire to use system deep login for workload

07:11.000 --> 07:14.000
at the station, so that it can detect.

07:14.000 --> 07:16.000
Oh, it's for me if you're running on the server.

07:16.000 --> 07:20.000
Or is there an alert manager's service running on the server?

07:20.000 --> 07:24.000
And the other plugin that I use is, it's renamed to TPM,

07:24.000 --> 07:28.000
but yeah, the bottom turtle plugin, which does the TPM at the station.

07:28.000 --> 07:31.000
So it registers the TPM with the CA.

07:31.000 --> 07:36.000
The CA makes sure that the TPM is legitimate,

07:36.000 --> 07:41.000
and uses that to bootstrap trust.

07:41.000 --> 07:46.000
So how does it work for a service like, for example,

07:46.000 --> 07:49.000
OpenBow to talk to Spire.

07:49.000 --> 07:54.000
So Spire exposes some UNIX domain sockets, which from which you can fetch certificates.

07:54.000 --> 07:59.000
So for example, we might have some policy set up that says,

07:59.000 --> 08:03.000
we want to issue certificates for OpenBow.

08:03.000 --> 08:07.000
So if OpenBow that service tries to talk to the Spire agent,

08:07.000 --> 08:09.000
it will retrieve a certificate.

08:09.000 --> 08:12.000
But if NJNX service tries to do the same thing,

08:12.000 --> 08:13.000
then we get an error vaccine.

08:13.000 --> 08:16.000
Oh, there's no certificate available for NJNX.

08:16.000 --> 08:19.000
It doesn't work.

08:20.000 --> 08:24.000
And yeah, then we just, to all our services that I'm running,

08:24.000 --> 08:25.000
oh my, where's my mouse?

08:25.000 --> 08:29.000
On my home lab, I just add, at the start of the service,

08:29.000 --> 08:32.000
this gets SVID snippet, which fetches this certificate

08:32.000 --> 08:33.000
for that specific service.

08:33.000 --> 08:35.000
So we have OpenBow fetch one.

08:35.000 --> 08:38.000
We have alert manager fetch one,

08:38.000 --> 08:41.000
with, for me, if you just fetch one.

08:41.000 --> 08:45.000
And then now that they're all certificates,

08:45.000 --> 08:50.000
I can, for example, configure OpenBow

08:50.000 --> 08:52.000
that it only trusts certificates.

08:52.000 --> 09:00.000
We can configure OpenBow to only allow us to fetch the web

09:00.000 --> 09:05.000
hook sequence from OpenBow, if it is the alert manager certificate

09:05.000 --> 09:06.000
talking to it.

09:06.000 --> 09:09.000
So now only my alert manager service can talk to OpenBow

09:09.000 --> 09:12.000
and fetch the web hook sequence.

09:12.000 --> 09:15.000
So yeah, with that, all our problems are solved.

09:15.000 --> 09:18.000
Mostly, everything talks to the alerts to each other now

09:18.000 --> 09:20.000
and stuff.

09:20.000 --> 09:25.000
And that's super cool, I think.

09:25.000 --> 09:27.000
And I want to talk a bit about, like,

09:27.000 --> 09:29.000
how does it actually work all this TPM stuff?

09:29.000 --> 09:31.000
Because it's kind of gloss over it.

09:31.000 --> 09:35.000
And yeah, the thing that's open that Spire does is,

09:35.000 --> 09:39.000
it does a complicated handshake with your trusted platform module

09:39.000 --> 09:43.000
on your device to make sure that this is,

09:43.000 --> 09:46.000
like, a legitimate TPM from a known manufacturer.

09:46.000 --> 09:51.000
Maybe I hard-coated the ID in some list in my home lab.

09:51.000 --> 09:54.000
And once you know that the TPM is,

09:54.000 --> 09:57.000
tell us the cool thing is that the TPM measures

09:57.000 --> 10:00.000
what software is being booted on your machine,

10:00.000 --> 10:04.000
like the firmwareable measure, what disk image is being loaded.

10:05.000 --> 10:08.000
I combine this with the fact that the NYXOS has support

10:08.000 --> 10:10.000
for this thing called measure boot.

10:10.000 --> 10:16.000
You can build disk images with NYXOS bundled as a EFI binary.

10:16.000 --> 10:20.000
And when your firmware loads your NYXOS image,

10:20.000 --> 10:24.000
it will load exactly all the hashes of the stuff

10:24.000 --> 10:25.000
that has started up.

10:25.000 --> 10:28.000
And you get some unique hash for the NYXOS configuration

10:28.000 --> 10:30.000
that you started.

10:31.000 --> 10:35.000
And yeah, we have functions for this in NYXOS,

10:35.000 --> 10:39.000
where you can create a little disk image with a DM verity

10:39.000 --> 10:40.000
protected boot-of-s.

10:40.000 --> 10:44.000
So the boot-of-s cannot be metled with.

10:44.000 --> 10:48.000
And this is all bundled in a work we call a UKI,

10:48.000 --> 10:51.000
which is like an EFI binary containing the kernel

10:51.000 --> 10:52.000
and the command line.

10:52.000 --> 10:57.000
And it points to the hash of our boot-of-s containing the NYXOS.

10:58.000 --> 11:01.000
So every time we have a new NYXOS generation,

11:01.000 --> 11:03.000
we have a different disk image,

11:03.000 --> 11:05.000
and it will produce a different hash in the TPM

11:05.000 --> 11:07.000
when the firmware loads it.

11:07.000 --> 11:09.000
And that kind of gives us a unique identity

11:09.000 --> 11:11.000
for each of our servers, right?

11:11.000 --> 11:13.000
Because each server will have a different NYXOS image.

11:13.000 --> 11:17.000
So it will produce different hashes when loaded by your BIOS.

11:20.000 --> 11:23.000
And yeah, building switch images with NYXOS is easy.

11:23.000 --> 11:24.000
We have modules for this.

11:24.000 --> 11:26.000
We have a report image builder.

11:26.000 --> 11:29.000
You can take an existing NYXOS configuration

11:29.000 --> 11:33.000
and turn it into a kind of immutable disk image.

11:35.000 --> 11:40.000
And yeah, then we can take such a disk image.

11:40.000 --> 11:42.000
Like we can take the UKI,

11:42.000 --> 11:46.000
and we can pre-predict what we would expect the TPM

11:46.000 --> 11:49.000
to see when a boot-of-this image is very meta.

11:49.000 --> 11:52.000
But we can have a NYX build that gives us exactly.

11:52.000 --> 11:55.000
Okay, this is the hash that the TPM will register

11:55.000 --> 11:57.000
if it boots it up.

11:57.000 --> 12:02.000
And then when our server registers with the Spire CA,

12:02.000 --> 12:05.000
we actually see, oh yeah, it registers with the same hash.

12:05.000 --> 12:09.000
So it can say like, you only get the premier-fios certificate

12:09.000 --> 12:11.000
if you actually boot it, the NYXOS config

12:11.000 --> 12:14.000
containing the premier-fios server definitions.

12:14.000 --> 12:17.000
And that way we get a super strong guarantee,

12:17.000 --> 12:20.000
like okay, I only want to give the TLS certificate

12:20.000 --> 12:22.000
of the premier-fios server to the premier-fios server,

12:22.000 --> 12:24.000
not to the alert manager server,

12:25.000 --> 12:27.000
and then by each thing in my home lab

12:27.000 --> 12:30.000
has a very strong unique identity.

12:33.000 --> 12:37.000
So now I'm going to try to do a life demo

12:37.000 --> 12:39.000
as a terrible idea.

12:39.000 --> 12:41.000
But we'll do it.

12:41.000 --> 12:43.000
And if it doesn't work,

12:43.000 --> 12:45.000
we can do questions instead.

12:45.000 --> 12:52.000
So I need to show you my screen once a second.

12:54.000 --> 12:58.000
Mirror.

12:58.000 --> 12:59.000
Yes.

12:59.000 --> 13:03.000
So I created a little NYXOS test that models my home lab

13:03.000 --> 13:05.000
and has all these servers.

13:05.000 --> 13:07.000
And when I hit test script,

13:07.000 --> 13:10.000
they're all going to boot up their images,

13:10.000 --> 13:12.000
register with the Spire server.

13:12.000 --> 13:14.000
Then it's going to open up my browser

13:14.000 --> 13:17.000
and show that everything can talk to each other, maybe.

13:18.000 --> 13:23.000
Wow, this loads.

13:23.000 --> 13:26.000
I'm just going to go to the next slide.

13:26.000 --> 13:29.000
Last talk, I didn't have links to source code.

13:29.000 --> 13:31.000
I fixed that now because it's for them.

13:31.000 --> 13:33.000
Everything of this open source,

13:33.000 --> 13:35.000
there's a NYXOS module for it.

13:35.000 --> 13:38.000
I had to contribute a bunch of things to Spire to fix it.

13:38.000 --> 13:40.000
And my home lab code is here as well.

13:42.000 --> 13:44.000
Wow, this is loading.

13:44.000 --> 13:49.000
Okay, metrics are starting up.

13:49.000 --> 13:52.000
Yeah, so each of these are building the images,

13:52.000 --> 13:54.000
booting the images,

13:54.000 --> 13:57.000
measuring it into virtual TPMs.

13:57.000 --> 14:01.000
And then it's succeeded.

14:01.000 --> 14:05.000
And it shoots its my own local CA,

14:05.000 --> 14:06.000
so we've got an error.

14:06.000 --> 14:09.000
So we have an alert manager instance.

14:09.000 --> 14:11.000
And we have a...

14:12.000 --> 14:14.000
For me, if he has instance, and you can see,

14:14.000 --> 14:16.000
for me, if he's scraping over TLS,

14:16.000 --> 14:19.000
and it successfully does it.

14:19.000 --> 14:22.000
So they trust each other, they trust each other's certificates.

14:22.000 --> 14:25.000
And we can see that my media server is down.

14:25.000 --> 14:29.000
So this should cause a alert in a few seconds

14:29.000 --> 14:31.000
that alerts should be able to be pushed

14:31.000 --> 14:34.000
to alert manager over TLS.

14:36.000 --> 14:38.000
To the do.

14:38.000 --> 14:42.000
I think I have a 30 seconds delay in the alert firing.

14:42.000 --> 14:44.000
Okay.

14:48.000 --> 14:50.000
Okay, alert fired.

14:50.000 --> 14:53.000
Showed up on the alert manager.

14:53.000 --> 14:55.000
And if there's internet,

14:55.000 --> 15:00.000
which I don't have, we should also see that the alert manager fetched the secret from OpenBow.

15:00.000 --> 15:03.000
And use it to push notification to my phone,

15:03.000 --> 15:05.000
but this part is not going to work,

15:05.000 --> 15:07.000
because the Wi-Fi is flaky.

15:07.000 --> 15:10.000
But yeah, end to end demo, you can check it out.

15:10.000 --> 15:12.000
And that was what I wanted to show.

15:12.000 --> 15:16.000
You can bootstrap your own cryptography bootstrap trust

15:16.000 --> 15:18.000
in your home lab without any secrets.

15:18.000 --> 15:20.000
Oh, and the others came in.

15:20.000 --> 15:22.000
That's...

