Introduction0:00
Allright.
Uh, good morning, everyone. I am Yedrick Kosinski, and this is—
Yeah, hello. I am, uh, known online as ComfyAnonymous, the original creator of ComfyUI.
And we are part of the ComfyOrg, the organization that, uh, is in charge of ComfyUI.
So, uh, I guess now that I have a mic, I'll ask again: who here has heard of ComfyUI?
Allright, allright. This half of the room: very knowledgeable, very nice, very nice. Um, for those unaware, we are an open-source, node-based design canvas intended for, uh, generative AI purposes, for, uh, multimodal, um, creative applications. We support image, video, audio, 3D, text, and more, uh, generative AI models.
Um, ComfyUI supports the absolute bleeding edge of generative AI tech. On day one, we have ComfyAnonymous here implementing it from not quite scratch, but it is redesigned from the original implementations. Uh, we offer open-source, locally hosted models that support NVIDIA AMD and Intel hardware, and we also support closed-source, API-accessible models that we, as the name suggests, just use an API to deliver to the user.
All of this functionality is also extendable with community-supported custom note packs, so anything we do not have the time to get to ourselves, the community does for us.
A big part of what makes ComfyUI special is the shareability of the workflows. Any image or video that was generated by ComfyUI has embedded metadata that lets you drag it back into the canvas and brings you the original workflow with all of the parameters that was used to generate it.
Uh, this sort of shareability and virality has really helped ComfyUI's traction. If you do a simple Google search on, uh, "ComfyUI workflows," you will find pages and pages and pages of results from the past year and a half, most of which are still compatible with modern ComfyUI versions.
Um, in terms of pure numbers, uh, this sort of shareability and virality has, over the past two years, taken us to the position of top 150 most popular GitHub repos of all time, with 78,000 stars.
Any comments, Comfy?
Uh, no.
Allright, allright.
Uh, with more of the traction numbers: we have 3 to 4 million active users, we have 20k daily downloads, we have 22,000 custom nodes made by 3,000 public developers that we enable in our ecosystem, and we've been adopted by Amazon, Apple, Tencent, Netflix, and more.
And pretty much any startup these days built around visual generative AI probably has ComfyUI working somewhere on their backend.
Um, why is ComfyUI popular? Um, it gives maximal control. You can go beyond prompts and interact with models that give you access to depth maps, line art, uh, masks, anything like that that is out in the space. If it's out, if it's open source, we probably either support it directly or the community has, uh, made it possible.
Uh, we are an all-in-one platform for both exploration for creatives and automation for developers. Sometimes those roles can also be switched, where developers want to explore tweaking models and seeing how things can be extended, and so we offer that as well through our custom node, uh, feature.
And because we are open source, we do not only depend on the output of the core team; we can trust the community to let us know anything they'd want us to work on, and also make anything we do not have the time to work on on our team.
Any comments, Comfy?
Uh, well, I think, uh, yeah. I think we haven't shown the, the interface yet, so.
Yes, we have only shown one screenshot of the interface at the very start. We will, uh, show that off as well. Don't do it n just do it later.
Yeah.
Allright. Story behind ComfyUI.
Origin Story4:43
Yeah.
Do you want to get into this, Comfy?
Yeah, the quick story is that, uh, basically ComfyUI started as my own personal project, and then I and then, uh, yeah, and then which, yeah, I started it in, uh, January, January 2023, and then six months later I was hired at Stability AI, so I spent one year at Stability AI.
They were using Comfy for, uh, for more, uh, like experimentation with, uh, internal experimentation with the models. And then I left Stability AI in, uh, June 2024, and then I joined up with, uh, Yolen and Robin, and, uh, we, uh, we made, uh, like the Comfy company, and, uh, yeah, that's, uh, things have been, uh, going pretty well since then, so.
Yeah.
Yep. And this picture was taken on the.
Yeah, we, we went, uh, that, that picture, we went on top of, uh, Mount Fuji, uh, which, uh, I don't recommend. It's, uh, ver-very, very difficult, but, uh, yeah, but we did it, so yeah.
I, I looked out, and my flight to Japan, uh, got rerouted to Alaska for 24 hours, so I landed in Tokyo five hours before they were going to be waking up to go to Mount Fuji, so I got to.
Yeah, so, yeah, so you, you missed the fun.
I missed the fun and then still got sick for a weekright afterwards.
Yeah. Okay. Yeah, so.
Allright. And back to now. ComfyOrg is indeed hiring. Uh, you can look at any opportunities on comfyorg/careers.
Yeah. Yeah, we yeah, we're hiring for a bunch of stuff, so if you're interested in, uh, joining us, if you're interested in open-source, uh, generative AI, well, that's, uh, maybe, uh, yeah, maybe we have a, a spot for you on our team, so.
Definitely check out the website. Uh, that is all for the official slides, but now is the fun part of showing the UI and taking any questions you may have.
Nice. I'm sure many on this side of the room who are familiar with ComfyUI know this standard, uh, Galaxy bottle workflow. Unfortunately, this spoils the results, so I'll just shake up the seed. Anyone not familiar with ComfyUI, this is all being locally rendered.
UI Demo7:09
Yeah, but this is a very old model. This is SD 1.5.
Yes, this is.
So that, that's why the results are not, not very good.
Yes, this model was, I think, the one that inspired your initial work on ComfyUI at the time.
Yeah, well, fine-tunes of this model. This is the base model, which isn't very good, but.
Yeah, this is.
But it's very fast, so.
It's very fast, but it is ancient tech at this point.
Yeah, it's almost three years old at this point.
Yeah.
So.
Ancient.
Yeah.
Allright. And there's the UI. Like, you're, you're, you're, you're looking at it.
Yeah, so basically what, uh, Comfy does for those who are not familiar, it kind of splits the whole diffusion pipeline into these different components. Like, a stable diffusion model is a diffusion model, a text encoder, and a VAE, which is why you have those three thingsright here.
So yeah, model, diffusion model, clip is the text encoder, VAE the VAE, and then so have the sampler, VAE decode, and save image, and that's, uh, that's basically a basic, uh, diffusion model pipeline. And what that lets you do is you can let me check which models you have on here.
Okay.
Not many.
Well, maybe SDXL is, uh.
Yeah, that one should work.
It should work for the.
I can, uh, type those numbers in for you.
Uh, you need.
In my, uh, rookie mistake of turning off my numlock.
Let's see how quick my disk drive is.
Yeah, this is all running on the laptop. That's why it's, uh, a bit slow, but, uh, once it samples, we'll get there.
Yeah, so, yeah, so this does look a lot better than, uh.
Yeah, this model is still also ancient. I think this one's.
Yeah.
Two years old, uh, at this point.
Yeah, this one's, uh, two years old.
Yep. We have more exciting workflows, though, if we browse the templates. If we want to go a little advanced, we've got there we go. This will not run because I do not have, like, 60 gigabytes' worth of models.
But here's what that workflow looks like.
Yeah, this is, uh, what a video workflow looks like, which you can see it's very similar from, uh, from one of the image workflows. It's just you still have the, the samp the same sampling node with different settings and same VAE decode node, which is kind of hidden here.
And yeah, so, so this is, this is the, the Wan 2.1 model. That's probably the best open video model at the moment. And you can see the, the pipeline is still very similar to even the first, uh, stable diffusion 1.5 model that was that we were presenting earlier.
So yeah, but what that lets you
do, the fact that you can, uh, you can go and change things. Like, say if I want to, uh,
like,
like this is, uh, one of, uh, a technique that, uh, this is basically, uh, this node where that I just added, what it does is it, um, it's a what I call a CFG trick, so it, it will add something to the, uh, to the sampling, to the CFG calculations of the sampling code.
So basically it's you can easily write these nodes, which, uh, which will change so you can go and just patch the pipeline this way, just by add just by either writing your own nodes or using nodes that already exist.
And for anyone familiar with CFG, it is a AI trick where you take the positive prompt, you sample on that, you take the negative prompt, in this case text and watermark, you sample on that, and with the magic of AI, you literally subtract the results from each other, and that, in some way, improves the image result.
Yeah, I think we can yeah, we I think we can take, uh, does anyone have any questions about anything? Like, uh, or anything in general related to ComfyUI?
We have a questionright here.
Model Basics12:55
So I see I can see on the screen something about clip text and code.
Yeah.
What's that relationship with the model that we're looking at?
Uh, yeah, those are basically those clip where you know what clip is?
No, it's a model.
Yeah, it's, uh, basically the diffusion models, they use the text encoder part of the clip model to
it's, uh, instead of so instead of passing the text directly to the model, they use this, uh, a text encoder because that way the model doesn't have to the diffusion model doesn't have to learn, like, all the to understand human language.
It can just learn the output embeddings of whatever text encoder you use. So yeah, so the, the this is, uh, basically the clip in ComfyUI represents the text encoder. The reason it's named clip is because before, like, on the stable diffusion models, they were only using clip as the text encoder, but in later models, it's more they started later models started using different text encoders that were not clip.
So the name I should yeah, the name should be changed, but, uh, yeah, so what this does is it essentially what this node does is it passes the text through the text encoder, and then the output would essentially be the output embeddings of or the last hidden the last hidden state, essentially, of the text encoder.
And that's usually well, it depends. It's slightly different for every for every model, but essentially it's the most of them, it's the last hidden state or the penultimate hidden state that is passed to the diffusion model.
So why there are two clip text encoders here?
Yeah, because this is a positive and negative prompt. Uh, this is how, uh, the CF like, because the models, how you sample most of these diffusion models is with, uh, a positive and a negative prompt, and that's using CFG, something called classifier-free guidance, CFG.
And what it basically the, the idea is that if you only sample with a positive prompt, so yeah, if I put CFG to one, that's essentially just sampling with a positive prompt. And you, you can see what happens when you, you only sample with a positive prompt.
It's, uh, you can see that the image is wait, this is worse than well, no, okay, it's because I have this node. Well, yeah, this is worse than it should be, but, uh, okay, yeah, if I yeah, if I sample with just, uh,
just a posi you see that it's, uh, the image is not very well defined. It's very chaotic if you only so what CFG does, it's a trick because if you think, uh, of all the possibilities of what the model can generate, it's kind if you it's kind of a way to push for, like, the CFG scale.
It does when sampling, it does positive minus negative prompt, and it's a way to push the sampling very more towards your positive and away from your negative. So the higher the scale, the more it will do that, which means you get a more defined image.
I don't know if, uh, my explanation makes sense, but, uh.
Yep.
Yeah.
Yep.
The VAE encoder is part of the model that diffuse stable diffusion?
Yeah, the VAE is because the what made stable diffusion be ex work extremely well.
For any function.
And, uh, yeah, what, what made stable diffusion be extremely popular is the fact that the, the image generation happens in compressed latent space. So instead of doing it in pixel space on a, like, let's say a 5000 a 512 times 512 image in pixel space, that's, uh, that's a lot of pixels.
Some earlier diffusion models did that, but they were pretty slow. Stable diffusion, it did this in, uh, latent space, which, uh, for a stable diffusion, the VAE is 8x compressed on every, uh, on every on the two, two dimensions.
So yeah, so instead of, uh, sampling a, uh, yeah, a 512 times 512, you would be sampling a 64 times 64 image, which is which is why these models are got so popular, because they were a lot more efficient than, uh, what came before.
So yeah, so that's what the VA the VAE is just a yeah, it's a VAE. It in-input is, uh, like 512 times 512 times 3 channel, and output would be, uh, would be, yeah, 64 times 64 times 4 channel in the case of, uh, of this model.
Awesome. Thank you so much. Thank you.
No problem.
So.
Allright, we have a questionright here. And I'll, I'll give you the mic.
Evaluation18:53
Thank you. So, um, ComfyUI is really in a lot of the examples is focused on the image generation as such, you know, kind of all kind of cool plugins. Um, I wonder if you have any good suggestions or ideas about evaluating the results, kind of like verifying or kind of like seeing this is good image or not a good image, uh, to, to kind of automate that workflow as well.
Uh, that's, uh, that's a difficult thing to do usually, because if, uh, it's the problem where, like, how do you define a good image? Because, uh, yeah, is there yeah, that there's some problems with, uh, because, uh, people's taste is very subjective.
So what is a good image for one person might not be a good image for another person. So yeah, it's, it's a problem they have. It's actually a big problem with the, like, user people who do who train these diffusion models, like user preference.
They, uh, when they when they actually add user preference data, their results get a bit worse because users like, uh, like the average user likes a certain type of image, which is not maybe might not be what, uh, what most what most people want.
So it's, uh, yeah, but, uh.
Any follow-ups?
No, it's more like I've seen kind of like critique models that you bring in, or.
Yeah.
You kind of have a prompt that looks at the image, uh, like a multimodal. But anyway, if there's nothing.
Yeah, there's a yeah, I yeah, we've had like, at least back when I was at Stability, we did have some, uh, we did experiment with some models that tried to just see, oh, like, get the output im output the image from the workflow, get some kind of rating from a model.
But, uh, it didn't work that well. So it's.
Okay, fair enough. Good question.
Allright, do we have any other questionsright now from anyone?
Raise your hand so I can see.
Gotcha. Do you have another one? Awesome.
So this is predominantly a workflow, and once you kind of like, uh, develop it, you do it in the UI. Um, any good tools around then, uh, running this more headless and kind of scaling this out and maybe building this into an app for kind of people using it?
Scaling21:30
Yeah, this is, uh, just, uh, yeah, this is one thing that, uh, because well, is this what ComfyUI is? It's actually you have this interface, but you also have a powerful backend behind it, which executes the workflows. Andright now there's, there's actually a lot of, uh, a lot of different inference service for these workflows, and eventually we'll be building our own.
So and yeah, and there's, there's already some, uh, a lot of, uh, third-party services that I saw that, uh, you can take your workflow, make an app out of it, and, uh, yeah, so you can already you can already do that, but, uh, just there's no just no official way of doing it, but there might, uh, there might be one in, in the future, so.
Okay, thanks for clarifying.
Thank you for the question.
Allright, any questions? Because we'll keep on talking about other stuff if there are no more questions, so be prepared.
API & Models23:07
Allrighty. Uh, one of the more recent additions to ComfyUI for a long time, we only supported open-source local models. In the past month, we've introduced API nodes, which for paid credits allow you to generate remotely.
Um, there is open up a template. We can do there we go. One of the models that recently came out was a, uh, Black Forest Labs, uh, context model, uh, currently not out for open-source usage in terms of being able to run locally, but they have made the APIs available.
Yeah, eventually they're supposed to release an open-source version, which, uh, well, we, we already support. They just haven't, haven't released it.
Yes, we are waiting for the green light.
Yeah.
And I would run this, but I have no internet connection, and that's one of the limitations of API nodes. You need to, you know, they're not ran locally.
Yeah, so yeah, I think there's some interesting
yeah, so we have, uh, yeah, yeah, we have a lot of different, uh, so the models that so that we support image, video, yeah, 3D. So we have a basic support for the, like, a Honeon 3D model, which is, uh, basically it's an interesting model.
It basically outputs a voxel type, uh,
yeah, like the, the 3D model, all these output is a kind of a voxel format, and then you and then so that's why in the workflow there's, uh, yeah, there's some, uh, code to so but the problem with these models, since it's kind of it generates some voxel format, and then you need to use an algorithm to convert it to mesh is that the mesh isn't very high quality, but it's still, uh, pretty impressive.
I did not have any of these models coming.
Yeah.
Uh,
so.
And we're currently, I guess, not well, I mean, we have local support for LLMs.
Well, there's a bunch of custom nodes with, uh, local LLM support. It's just not a core Comfy thing yet. It's just we're more focused on, uh, on, like, image and video and all these, uh, more visual. Oh, we also support audio, an audio model now.
So it's, uh, yeah, it's not as good as some of the, uh, proprietary models out there, but it's, uh, yeah, it's pretty fun to play with.
And there were some more, I think, uh, audio models that came out this week.
Yeah, but those, those are text-to-speech models.
Gotcha.
Yeah, those which we, we may support. We'll, we'll have to see if, uh, because they're already supported as custom nodes, but, uh, yeah, before to yeah, it's just to integrate them in core Comfy, there needs to be, like, a reason to, like, if, uh, give them some extra control or some extra, like, extra knobs to turn, or else there's not much point.
Yeah. I'm interested in any questions from this side of the room that maybe wasn't too familiar with ComfyUI at the start. Uh, do you have any questions, comments, inquiries? Allright. I will hand you the mic.
Yeah.
Use Cases27:11
Uh, sorry, it's me again. So does ComfyUI have a use case for the virtual try-on where, you know, we upload the image of the model, uh, either mannequin and the garment, the clothes, so that it generate the virtual try-on images?
Yeah, like, for example, the new Flux context model can, can do that, uh, I think. So yeah, there's a few different there's some open-source ways, and there's some, uh, some ways using, uh, the API nodes. But, uh, yeah, virtual try-ons, it's something that seems very popular, so there are, uh, there are a bunch of workflows for it.
Okay, so we can find it on the ComfyUI and try it out.
Uh, yeah. Yeah, if you if you search, you can find, uh, you can probably easily find a workflow for it. The only thing you might, uh, it's just some of the it's just that the field evolves so fast that, uh, sometimes, uh, workflows you find might be slightly outdated.
So but if I was doing that, I would first try the new Flux context model, since that seems to be, uh, the best one for that. But, uh, yeah.
Uh, the name is new Comfy what's the model name? New.
Uh, Flux Concept.
Flux Concept.
Context, yeah, I keep okay, yeah, Flux sorry, Flux Context.
Context with a K.
Yeah, Context with a K.
Context.
So.
Okay, thank you. Thank you so much.
And to also follow up on that.
Yeah, andright now yeah,right now it's an API node only, but they're they should release the, uh, the open-source version soon. So yeah, so once that's once that's released, you'll be able to run it on your on your on your machine with the ComfyUI.
Yeah, to follow up on virtual try-on, this is actually something that people have made workflows in the past year. When we were in Japan, when we had a meet-and-greet there, there were some people who actually made workflows specifically for that.
Back then, there weren't some of the models like Context now are very good at a, "Hey, change this one thing." At the time, there weren't. So the workflows you'd find probably have a few dozen nodes, basically finding using one model to find the masks of, like, what to change, then another model to inpaint those masks of the actual thing you want to change.
Now the models are a bit more, uh, advanced, where you can just say, "Hey, I want to edit this," and it does it. And you, of course, combine up the masks as well in case the model gets a little, uh, a little rowdy and tries to change things you don't want.
You can always add masks to keep it contained.
Allright, any more questions on this side of the room?
Uh, it's a good question.
So sorry, I joined the session very late, but, um, if we want to generate any kind of image, I think this allows us to write a prompt, and then it allows us to generate image. Is that correct?
Yes.
Okay, so for example, if you want to have a tool that automate building multiple images based on, let's say, character, like, if I if I want to have defined a character, and if I want to generate a stories based on the characters, does this allow it?
Yes, well, yeah, what you need is, uh, there's a few different ways to because I assume, yeah, you want to generate a consistent character.
Yes.
Because depending on what you want, you can either, uh, train a LoRA for your character or use one of the newer model, like, uh, like the Flux context model. Like these, uh, like, very recently, there's all these, uh, edit models that have what I call edit models, which are basically, uh, they got very inspired what, what, uh, 4.0 was doing.
So.
Which one do you suggest?
What?
Which one do you suggest?
Uh,right now the best one is, uh, the Flux, uh, the yeah, the Flux, uh, context model. But, uh, like I said, it's onlyright now it's only available through an API, and but, uh, should be open-source soon. And there's some other ones too, but, uh, that one, yeah, what you can do with, uh, with the, the context is just, uh, like, some you give it a reference image of a character, and you say, "Oh, make that character do this," and it actually keeps the character consistency extremely well.
Just, uh.
So there is there is a way to, uh, maintain character throughout the story generation,right?
Yes. Yeah, well, what you would do is you would have, uh, yeah, first you generate a your character of an image of your character that you're happy with, and then you would, uh, you would pass it to this model and say, "Oh, put this character in this scene, put this character in that scene," and then you generate your image is based on this reference image of the character.
Okay, thank you.
And to follow up on that, one of the advantages of a node-based system is with the way that is set up, all you can currently edit in it are some of the parameters and the text prompts, but you could also apply the LoRAs.
LoRAs are, uh, low-rank adaptations to the model. Um, and because it's node-based, you can also mask the specific area each of those low-rank adaptations would apply to. So let's say you have two LoRAs trained, one for character A, one for character B.
Uh, what our node-based system allows is to say, "Hey, in this area of the image, I'd like this LoRA to be active. Maybe at this strength." You could even schedule it in terms of that. And the other area of an image, you can have, "Oh, I want this other character LoRA to be active."
So if you even if you, uh, if all-in-one model, like Context, doesn't quite do what you want, there are multiple ways you can sort of coerce these models to kind of do it. With a basic, uh, prompt-based system, there are, of course, limitations, but because we are node-based, you can do, you know, there's two things for the prompts there.
You could set that up to be 10 nodes, and some of those nodes apply a specific LoRA to a particular image sorry, to a particular area of an image.
Do you also, uh, do you.
Do you also recommend LoRA or, uh, the other one?
Um, if you don't have, uh, like, much experience in this space, I'd recommend the context model, mainly because you just you just have to type in the prompt, and it does the work for you. The other one, especially back before these sort of, you know, edited via text models existed, was sort of the brute-force way of getting what you want.
But you could really get what you want because you could train it on anything you want. The models don't have to be aware of what it is. And the only disadvantage is you need to have enough training images, so, like, between 10 to 30, to actually get your subject to appear the way you want them to.
With these newer edit models, you only need to give it one image.
Thank you.
No problem.
Yeah.
Allright, any questions here?
Or back on that area of the room? I can walk.
Allright, Comfy, what do you want to talk about next?
Uh, well, yeah. Yeah, well, we since we mentioned LoRAs, like, LoRAs are one of the basically, what they what they are is a, a patch on I call them a yeah, they're basically a patch on the model weights, which is, uh, or a more efficient way to train a concept or multiple concept in, in a in a model.
Advanced Techniques35:23
And yeah,right now we don't it's basically just if you want to train a model, instead of training the full model, you would train this small patch on the model. And this allows you to well, you can train styles, specific characters, anything.
So yeah.
Yeah, we can we can skip showing it off. This was for the, uh, Japan presentation where this.
Oh, interesting.
This LoRA is for, for an anime character. That goes hard in Japan. Probably doesn't go very hard at a AI conference.
Oh.
So.
But this is how you would do it. You would just chain the model there.
And these.
Examples running in, like, the devices.
Uh, what do I have any SDXL?
Which model is this?
I don't know. Okay, this is for 1.5.
Yes.
Okay.
And, uh, they're for Japan.
Allright, well, we can still show them off, but, uh.
Allright, we can try.
Yeah.
Oh, and these would probably look very poorly on these models, but we can give it a shot.
Well, I use the anime one, but.
Okay, we can use an anime one.
Okay, that's the anime one.
Yeah, it works.
So yeah, just, uh, I mean, if you try that prompt, it's, well, probably not going to.
Yeah, we can, uh, we can do that in a bit. Allright.
Okay.
What else would you like to talk about?
Uh, yeah, well, you can just try. See? Yeah, well, we can press run and.
I don't know if we should. I don't know if we should press run.
Okay. Yeah, you're just, well, okay.
Yeah, we can skip. Yeah, this is, uh, assignment to do at home, I suppose.
Uh, well, yeah.
But we have other models that we support. Let's see. Yeah, apologies that we do not have much live demos. Uh, uh, there were some setup last minute in terms of us attending the conference, so.
Yeah.
But we are here.
Sorry about that.
Uh, here are just some control net examples where we can't show the inputs, but we can actually.
Yeah.
I guess we can we can trust the template system to kind of show what that's about.
Yeah, control nets are just one of the many ways to, like, have more control of these, uh, of the, the models.
Yeah, so the examples here would be the inputs that were used to actually generate these images.
Yeah, but those might be like, control nets might no longer be very useful because now there's all these edit models that are coming out. So yeah, it just means that the space is, uh, is evolving. But, uh.
Ah, so here's a more advanced workflow where it applies, I believe, different prompts to different areas of the image.
Yeah, this is, uh, different prompts to different areas.
Yeah, we can actually make this one go on the default SD 1.5 model. That one will, will work.
Hmm.
Okay. Ah, yes, this is the old way of prompting things when the models kind of, yeah, to really coerce them.
We will fix the seed. Okay, and let's see how the laptop handles this.
Yes.
See, assuming there's no loaded images, this should just work.
Yeah, at least half the workflow should work.
Yeah.
So.
Yeah, this is a very old workflow, but, uh, I think it still works on even the most recent models.
Uh,
yeah.
And we can, uh, we can change the prompts maybe it's more obvious, but I believe the prompts are basically doing a different time of day on some of these.
Yeah, yeah, it's basically different time of day on, like, if you go. Yeah, top is, like, night, and bottom is daytime.
Yeah.
Yeah, just, uh, yeah, so this is just one of many ways you can get, like, more control. This is just a way of applying different prompts in different areas of the image. And like I said, I think it, it still works even on the most recent models.
Yep, yeah, everything that basically started from the foundation, uh, Comfy set up two years ago.
Yeah.
Most of those any of those tricks or applications still apply to newer models.
Yeah, yeah, because they're general, like, diffusion model tricks, and we're still using diffusion. So yeah, yeah, so that's what makes Comfy nice is that if once if a new diffusion model is implemented, usually you can use all the old tricks if you want.
Some of them might not be useful anymore, but you can still use them.
Yeah, the models have also gotten bigger and harder to run locally in some cases on some hardware. So, uh, some of these tricks would, you know, make things run quite a bit slower. Um, in the early days of image generation, a lot of the improvements were with community fine-tunes who would take, you know, vast data sets and improve the base model.
You may have noticed I was a little nervous running a model, uh, a few minutes ago. The reason for that, that was one of those fine that was one of the sort of days of back of community fine-tunes.
Uh, the data sets they used may not always produce, uh, the most, uh, conference-friendly content.
Yeah, yeah, there's, there's some interesting things that happen when a model is slightly broken because since it's a diffusion model, oh, if it's slightly broken and you're generating a, like, a character, the first, like, the first step might produce a, like, a skin color blob, which means it might converge to, to a naked person, basically.
So yeah.
Yeah, and given there are community fine-tunes that basically everyone trusted to produce better quality images, those are usually generations that you first review and then show rather than press queue and then, uh.
Yeah, but that's the power of, uh, running things locally. You don't have, uh, any problem. Uh, you can do whatever you want. So yeah.
Yeah, with newer models and bigger ones, the training sets are a bit more constrained. So you have.
Yeah.
The pros and cons of that.
Well, it's just they're, they're better. They make less, uh, random mistakes.
Yeah, you can be more you can trust more that when you put in a specific prompt, it will not hallucinate as much.
Hmm.
Yeah.
Final Q&A43:28
Allright, so in terms of we mentioned that we are hiring. I believe we're looking for positions on.
Well, everything, pretty much.
Yeah, everything. Back end, front end.
Yeah, core.
Cloud deployment.
Core inference, uh, yeah, cloud. Yeah, just yeah, go look at our careers page and, uh.
Yeah, it's Comfy.org/careers.
Yeah, and if you haven't tried the software, go try it. You can just if you all you need is a, a decent GPU, and you can run it locally. Or you can use the API nodes and yeah.
Yeah, people have gotten some of the early models to work on extremely old GPUs, like.
Yeah.
80-year-old GPUs.
Yeah, yeah, one of the strengths of Comfy is that pretty much any hardware well, any NVIDIA hardware, the model will usually run. It might be extremely slow, but it will usually run. So yeah, so are there any final questions?
Yep,right there.
I can screen-read it.
Uh, hold up, I'll give you the mic. There's a process to this thing.
Thank you.
Yeah.
Um, I've tried using Comfy, and I was just wondering, like, if you could give us, like, a quick synopsis of what do you think about Comfy versus the alternatives that exist? Like, why would you sort of say Comfy is the one that people should start with or stick to?
I have no idea, like, of the depth of it, so just give me, like, a seminar of that, please.
Uh, Comfy is, uh, you should use it because it's the it's the mo basically, it's the most powerful one. So if you, uh, like, everyone who like, it's basically the, the end game for, for these, these types of interfaces.
So there's nothing that gives you more control, that has more community support, that has more extensions. So the only downside it hasright now is it's, uh, it's a bit, uh, difficult to get into, but, uh, we are working on that.
So.
Thank you.
Yeah, node-based systems, especially if you're not used to them at first, can be quite intimidating. And as, uh, Comfy mentioned, one of the greatest assets of ComfyUI is that it is community extendable, and it is open source in that anything that the core team may not be able to get to, there probably exists a community solution for that or to do something.
Like, like, we mentioned the slide, there are, I believe, 22,000 custom nodes within, like, 3,000 note packs made by, you know, 3,000 separate developers who are all passionate. Uh, if you go to other places, you will not always have, you know, the certainty as, oh, can I run this locally?
Do I know all my data's safe? And if you are in a, for example, an enterprise setting, data security might be a big thing to avoid becoming the next headline in terms of a data leak or a ransomware attack.
So being, being able to actually look at the source code, if that's your thing, or having your team be able to look at the source code, you can contribute any fixes, uh, in terms of optimization and performance. We are pretty much state-of-the-art.
Uh, Comfy over there, when the new model comes out and he hears that there is a way to run it faster, he implements it. Or one of us on the team implements it. So.
Is there any recommendation you'd have, like, getting support? Like, where would you, like, is there a Discord channel?
Um, there's a Discord channel that we have for Comfy org. We also, as we post the slides, if you just Google ComfyUI, there will be most likely thousands of YouTube videos. Um, there's even some people who have taken, uh, they've seen the opportunity of the difficulty of ComfyUI, um, and they are, for example, having paid, uh, tutoring classes for it, which is a bit of a eye-opener for us because that says we should probably do a better job onboarding the users if, uh, people are, you know, making money that way.
But there should be a lot of resources out there for you.
Allright, any other questions? Allright, over there.
Yeah, thank you. Is there currently a published product roadmap?
Uh, if you mean, uh, like, what we are currently, um, well, we haven't started, really started, actually started yet, but eventually we'll have a solution to run these workflows in the cloud. And, uh, yeah, how exactly it's going to work because we the, the thing is, before doing that, we want to fix there's a, a few issues we have to fix, like the for example, we want to make installing and dealing with the custom node that you install.
We want to make that a lot smoother, make, uh, the interface better, add a yeah, what, what we're going to do is, uh, improve the interface key well, the, there's always going to be the node interface, but, uh, we are most likely going to add another layer on top of it where you can have a more build a more traditional interface out of your workflow graph.
And that will fit in with, uh, well, with the cloud stuff that we're going to be doing eventually. So yeah, so that's, that's the direction where we're going in. But, uh, the thing is, in this space is that things change a lot.
So a new model that comes out tomorrow might, uh, might mean we need to, uh, pivot a bit. So, uh, that's why I'm not, uh, I'm not giving any promises. So yeah.
Yeah, because, like, the, the first thing that went to my mind is we had the gentleman ask a question about can we serve these workflows up. So it's like, if you can access a workflow through an API, you can have, like, a single power user building out massive templates.
Yeah.
That maintain, like, style and brand guidelines or, or story or character or design. And then if you have, like, role-based access control, you could have, like, just a general user in there saying, "Hey, I need to generate this workflow based on these parameters.
I can't touch anything else in there." It's like, is that is, like, being more enterprise or team ready?
Yeah, like, this is one of this is a direction we are going into. So, like, having, uh, just the, the basics for that would be first a good cloud inference service where you can run workflows very well and have all the custom nodes work.
And once we solve that, then all that other all that other stuff becomes a lot easier. So yeah.
Thank you.
Yep, and to follow up on that, at the end of the week, we will have a blog post about some of the things we are working on. Um, for the because we are planning to allow cloud services, but first, as he as, uh, Comfy said, we need to work out dependency issues.
So we'll have a bunch of features being announced there. For example, we'll have a subgraph option where you can combine a bunch of nodes, put it into one node, and you can double-click into it as, like, a separate workflow.
Uh, solving dependency issues, we are at custom nodesright now. We can request different Python packages, making sure all of those could get either properly isolated or have more ways for them to report their compatibility. Because once local becomes much better to run, that means our life trying to get this as a cloud product will also become smoother.
Yeah, I think, yeah, we are out of time now. So.
Yep.
I would, uh, would like to thank everyone for coming. We, uh, yeah, and I hope you, uh, you learned something.
Yep, yeah, thank you for all the questions. Greatly appreciated.
Yeah.





