Repository navigation
question about running many queues for webservers and workers #96
Description
Activity
Spot-on use case and analysis of the situation. You're right to be concerned, but fortunately we can mostly keep it under control.
The Redis server allows something like
min(system_max_file_descriptors, 10000)connections, so we can get away with this one-connection-per-queue pattern for small deployments with lots of queues or large deployments with few queues, but otherwise we can use them up pretty quickly; if we have a medium deployment of 20 servers, 4 processes/server, 30 queues/process, and 2 connections/queue (typical case), we're halfway to our limit!Quick breakdown: queues can have 3 kinds of Redis clients:
- general client used for most stuff
- event client
- used for pub/sub - mostly for producers finding out their job is finished
- not used if
getEventsandactivateDelayedJobsare false (and latter is false by default)
- blocking client
- used by workers waiting for a job to come in (
brpoplpush) - not used if
isWorkeris false
- used by workers waiting for a job to come in (
Every queue always has its own event and/or blocking client if necessary, but they can reuse/share the general client (but don't by default). This means that if you have 30 producer queues and none of them need to receive completion events, you can use just one connection for all of them, but if you have 30 workers you will need to use 1 general connection plus 30 blocking clients. Just make sure you also have the proper settings flags set depending on the role of each Queue instance to avoid these secondary connections when you don't need them - it looks like you've already got that sorted on the producers in your example, but you should be able to do
getEvents: falseon the workers as well.To share the general client, you can pass a node_redis
RedisClientinstance as theredisargument of the Queue settings; make one, pass it to all of your queues. You can see the implementation here and we briefly mention this in the queue config docs, although I just noticed the explanation there kind of trails off - would definitely accept a PR to improve that.So your above worker example could look something like:
const redis = require('redis') const Queue = require('bee-queue') const sharedConfig = { getEvents: false, redis: redis.createClient(process.env.REDIS_URL) } const emailQueue = new Queue('EMAIL_DELIVERY', sharedConfig) const facebookUpdateQueue = new Queue('FACEBOOK_UPDATE', sharedConfig) emailQueue.process((job) => { return Email.deliver(job.data) }) facebookUpdate.process((job) => { return FacebookUpdater.process(job.data) })
Fully implemented, if our original 20 servers were half producers and half workers and we don't use
getEvents, we can have just 40 total producer clients, and 40 * 31 worker clients - around a quarter of where we started. Of course, I doubt all 40 worker processes need to process on all 30 job types, so you could probably reduce it much further.@bradvogel I recall you mentioned that you run a pretty good number of separate queues, but I'm not sure how split up those are between services or how many total connections your Redis instances typically have open. Have you paid much attention to total connections by any one server/process or to total redis connection count, or taken any specific measures to help keep it under control?
Reacted by Dylan Jhaveri, Chia Yuan Chuang, Kaushik Thirthappa, Chandan and Evan Adams@LewisJEllis Thank you for your thoughtful and thorough reply, I had a feeling something like this would be possible.
I'll give this a try today and see how it goes. I'll happily take a stab at improving the docs, specifically related to the shared RedisClient
@LewisJEllis thanks, I've tested this out, seems to be working as expected. I submitted a PR to add some clarifying documentation and an example to the README.
Please let me know if anything in there is inaccurate or could be explained in a better way. Thank you!
I recall you mentioned that you run a pretty good number of separate queues, but I'm not sure how split up those are between services or how many total connections your Redis instances typically have open.
We have some redis clusters with upwards of 5000 open connections.
Have you paid much attention to total connections by any one server/process or to total redis connection count, or taken any specific measures to help keep it under control?
In the past we've consolidated to common redis command clients (as opposed to event/blocking connections), but haven't made it a high priority.
So basically yes this is the way we've handled many connections internally, and what we'd recommend for others.
Reacted by Dylan JhaveriClosing as this issue seems resolved (but useful for historical context 😄). Feel free to reopen.
Reacted by Dylan JhaveriIn the past we've consolidated to common redis command clients (as opposed to event/blocking connections), but haven't made it a high priority.
@skeggse I'm not sure what you meant by consolidated to common redis command clients when we have a blocking client (aka a Worker). Can you help me understand how would I make the worker do the job scheduled without the blocking client?
Reacted by alam38Hey, I wanted to ask the same question that Gask was asking above.
In the past we've consolidated to common redis command clients (as opposed to event/blocking connections), but haven't made it a high priority.
I'm not sure what you meant by consolidated to common redis command clients when we have a blocking client (aka a Worker). Can you help me understand how would I make the worker do the job scheduled without the blocking client?
For further context on my specific use cases. I have multiple docker containers running the same worker queues. Any help consolidating the number of redis connections would be greatly appreciated.
Adding @LewisJEllis for extra visibility. Please let me know if any extra context would help.
This seems to be a common pattern with background workers in a web application, but I haven't seen it explicitly documented so I want to open this up as a question. If necessary, I'm happy to provide a PR with a documentation update.
Let's say I have a web app and I have 30 different background jobs that get run. For example:
The way bee-queue (and bull) are set up, each of these 30 background jobs would have their own "queue". Below is an example showing the first two queues.
As one might imagine, with 30 different background jobs I will have 30 different instances of
Queueon each webserver and 30 different instances ofQueueon each worker server.From the docs:
My Questions
Webserver
Worker server