Visitar URL original
serial: TCP socket backend for serial and console by weltling · Pull Request #8874 · cloud-hypervisor/cloud-hypervisor · GitHub
Skip to content

serial: TCP socket backend for serial and console - #8874

Open
weltling wants to merge 6 commits into
cloud-hypervisor:mainfrom
weltling:serial-tcp-backend
Open

weltling wants to merge 6 commits into
cloud-hypervisor:mainfrom
weltling:serial-tcp-backend

Conversation

@weltling

Copy link
Copy Markdown
Member

Add a TCP backend to the SocketConsole so the serial port and the virtio console can be exposed over TCP.

The tcp=<host>:<port> option is accepted on both --serial and --console, following the QEMU chardev socket roles:

  • server=on binds the address and serves one client at a time, buffering output until a client connects.
  • server=off dials a remote listener, and reconnect=<secs> redials on that interval after a disconnect.
  • wait=on holds boot back until a client connects, valid only for a listening server.

This completes the second part of the linked issue and builds on the Unix socket unification.

Closes #8633

@weltling
weltling requested a review from a team as a code owner September 13, 2026 22:03

@phip1611 phip1611 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No detaileled review yet but LGTM on a first glance! This is something we need as well for sure.

What about #8633 (comment), using universal FDs? No need to implement it here as well but what are your thoughts?

Comment thread cloud-hypervisor/tests/integration.rs
_test_socket_interaction(ConsoleKind::Console);
}

fn _test_tcp_interaction(kind: ConsoleKind) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

could you please also add an live migration test? I suspect currently it will fail. In our fork (different impl, but same intend) we also had to change the socket path on the destination.

https://github.com/cyberus-technology/cloud-hypervisor/blob/616807da1470ce063abca2ab00ee5482f72c7329/vmm/src/api/mod.rs#L296

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point about migration. This series already handles client mode console migration, so I added tests for it.

The server role is a deeper topic and not specific to TCP. The Unix socket console is also same in terms of only supporting the listener mode. Supporting server=on for migration would seem a separate effort.

Thanks

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I am not mistaken you currently expect the socket to be available on the same name/path on the destination without a possibility to change paths - unlike what we did in our PR (link above) - is this correct?

It might be okay that we do not need this and that in our setup the paths are always equal. However, I have to double check that.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, that's correct. There is a dependcy on the migrated config and there is no mechanism to change it today. In the listening role, it'll rebind the same address. To be more flexible here, it'll need some changes to the mechanism that carries the TCP configuration. And the same applies to the UNIX socket path, probably.

Thanks

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let me sync with @hertrste and get back to you. Thanks for the information!

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to change the URL on the new destination. We can do that in a follow-up. Thanks for doing the groundwork here!

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Gotcha, thanks for double checking. Yep, I would also suggest considering migration out of scope of this PR. In particular, there seems to be a general dependency on the config that I believe might be a good point to think about. I've been checking the equivalent functionality in QEMU and the difference here is that the particular piece is kept outside of the migratable configuration, so then it is flexible and allows better handling even for the server=on case.

@weltling

Copy link
Copy Markdown
Member Author

No detaileled review yet but LGTM on a first glance! This is something we need as well for sure.

What about #8633 (comment), using universal FDs? No need to implement it here as well but what are your thoughts?

Thanks for the quick check!

The universal fd direction - I recall it has already been discussed in a broader way concerning all backends. Seems more like a global design that has to be decided first.

Thanks

@rbradford

Copy link
Copy Markdown
Member

@phip1611 Are you happy with @weltling's changes? You still have "Request changes" review

@phip1611 phip1611 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is close, just a few cosmetic changes for better/simpler/more readable code!

Comment thread virtio-devices/src/console.rs
Comment thread virtio-devices/src/console.rs Outdated
Comment thread vmm/src/config.rs Outdated
Comment thread vmm/src/serial_manager.rs Outdated
Comment thread vmm/src/serial_manager.rs Outdated
Comment thread vmm/src/vm_config.rs
Comment thread cloud-hypervisor/tests/integration.rs Outdated
@weltling

Copy link
Copy Markdown
Member Author

This is close, just a few cosmetic changes for better/simpler/more readable code!

Thanks for the second look. Ready for the next round.

Thanks

@phip1611 phip1611 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks 🚀

@rbradford rbradford left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm concerned about the connecting outwards. At this point our seccomp filtering is getting kinda useless (sure, it doesn't yet contain exec)...although I guess I already lost that battle when I let TCP support land since that thread meant that vmm thread also needed it.

ClientStream::Tcp(stream) => stream.write(buf),
}
}
fn flush(&mut self) -> io::Result<()> {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: whitespace line here.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed, thanks.

Comment on lines +26 to +29
enum ClientStream {
Unix(UnixStream),
Tcp(TcpStream),
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's annoying we have to do this but that's fine :-)

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, that's the boilerplate. A trait object cannot return a cloned stream, since it returns Self.


// Connect and register the client fd for input, leaving failures for the
// timeout to retry.
fn tcp_connect(&mut self, helper: &mut EpollHelper) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel like this functionality of connecting out is a bit scary? It could easily block VM operations. Now I know we do something similar for vhost-user devices already. But are we sure we need both client and server support?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, both are reasonable concerns. I have the considerations below.

The dial phase blocks only for the initial connect, right after the socket becomes non blocking. An additional measure can be to bound it with the connection timeout, so then an unavailable remote cannot stall the console thread.

With the client/server roles - IMO both have their place. Kernel debugging over serial is a concrete case, both Windows and Linux (and MSHV debugging is the same scenario). Removing an extra hop like socat is a good simplification. The reverse case for client would be, when Cloud Hypervisor cannot accept inbound connections behind NAT, plus the QEMU interoperability.

Thanks

Comment thread vmm/src/config.rs Outdated
.map_err(&map_err)?
.unwrap_or(Toggle(false))
.0;
let reconnect = parser.convert::<u64>("reconnect").map_err(&map_err)?;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we have a sensible default value for this?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, totally makes sense. It'll make the client mode actually work without extra flags. I've added a default of 1 second for now, please let me know otherwise.

Thanks

}

// A client mode TCP console dials over IPv4 or IPv6.
fn create_console_socket_seccomp_rule() -> Vec<SeccompRule> {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This feels like a really big hole to punch in our security defenses.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this is the first device thread with AF_INET/AF_INET6 socket and connect. So it is real outbound trafic, unlike the AF_UNIX only rules on the vsock and vhost threads. socket is already restricted to AF_INET/AF_INET6, and connect cannot be filtered by destination since the address is behind a pointer. To contain it, I can add these syscalls only when a TCP client console is configured, so every other console thread stays as locked down as today.

Thanks

@weltling

Copy link
Copy Markdown
Member Author

I'm concerned about the connecting outwards. At this point our seccomp filtering is getting kinda useless (sure, it doesn't yet contain exec)...although I guess I already lost that battle when I let TCP support land since that thread meant that vmm thread also needed it.

Thanks for looking into this.

Another observation - QEMU runs the chardev socket backend in the main process and dials from the main loop, so even with -sandbox on the serial path is not blocked from the network. socket and connect are available process wide. Cloud Hypervisor has per thread seccomp and is already stricter than that. But it is also worth noting QEMU is not considered less secure for allowing this process wide.

Thanks

The console buffer served a device byte stream to a single Unix socket
client. Extend it to serve a TCP client too, reusing the same buffering
and single client reconnect for both transports.

Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add a TCP console mode next to the file and Unix socket modes, on both
the serial and console ports. A server binds and listens for a client
while a client dials a remote and reconnects, and wait holds boot back
until a client connects. The serial manager and the virtio console both
serve the byte stream over TCP, reusing the buffering and single client
reconnect shared with the Unix socket console.

Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
A client mode TCP console opens and configures a socket from the serial
manager and virtio console threads, so widen their seccomp filters to
permit it, restricted to IPv4 and IPv6.

Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Exercise the TCP console over loopback for both the serial and the
virtio console frontend, connecting a client and checking the byte
stream in both directions.

Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Describe the TCP backend for the serial port and the virtio console,
covering the server and client roles and the wait and reconnect options.

Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Migrate a guest where serial or virtio console is a client mode TCP
socket and check the console still carries output on the destination
after it redials the listener.

Assisted-by: Claude:Opus-4.8
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

serial: Support TCP socket backend for serial and console

3 participants