Generative AI Update: Comparing ChatGPT, Claude, and Gemini while Researching the Metaverse Characteristics of Social VR Platforms

NOTICE: In this blogpost, I go into sometimes great detail about how these three generative AI tools work, comparing them in two ways:

– comparing how these tools work with the exact same text prompt; and
– comparing how they worked in August 2025 versus February 2026.

There’s an executive summary (Section 4) at the very bottom of this long, loooong blog post if you just want to skip to the highlights, and ranking.

If you need an introduction or a refresher, you might want to read this blogpost first: An Introduction to Artificial Intelligence in General, and Generative AI in Particular, which includes slides from lectures I gave on the topic in November and December of 2025.

SECTION 1: Introduction

In his 2024 book Co-Intelligence (still my go-to layperson’s guide to generative AI), Ethan Mollick says that one of the best ways to determine how well a particular generative AI tool works is to ask it questions about a subject that you already are an expert in. Why? Because it will be much easier for you, the human expert in the topic, to find errors and hallucinations in the answers.

Since last summer, I have been typing the exact same prompt into the “big three” general-purpose GenAI tools Ethan recommends: OpenAI’s ChatGPT, Anthropic’s Claude, and Google Gemini. I have been meaning to write a blogpost about my experiences with this first round of testing since September, but I have been too occupied with my paying job as an academic librarian to find an opportunity to do so—until now. (Please note that I have been using an em-dash, correctly, for many years before generative AI came along!)

So, today I decided to redo my original text prompt, using the latest versions of these three GenAI tools as outlined by Ethan in the latest edition of his AI Guide, which has posted to his Substack newsletter on Feb. 17th, 2026 (here’s a link).

I consider his advice to be quite valuable, as he seems to spend a lot of time working with the most popular and powerful GenAI tools, and keeping on top of the changes and advances in the technology. In this newest edition of his AI Guide, he discusses the shift from chatbots (where you have a conversation with the tool) to agents (where you give a specific, defined task with instructions to the tool, and it goes away and does the task and returns with results).

In all cases, the initial text prompt is the following:

What are some characteristics common to all metaverse platforms? How do these characteristics apply to social VR platforms? Please give me a chart comparing these characteristics for the most popular social VR platforms.

Please note that I have deliberately given the task of defining “popular,” and picking the social VR platforms, over to the generative AI tool (and I got some rather interesting results back!). Because I consider myself an expert on social VR and the metaverse, I should be able to spot inaccuracies, errors, or outright hallucinations in the responses I get back from these GenAI tools. In the next section (section 2), I compare and contrast the results I received from the above text prompt from:

  • Claude by Anthropic
  • ChatGPT by OpenAI
  • Gemini by Google

All three of these tools come with different versions. In all cases, I will use the most powerful version recommended by Ethan Mollick in his latest AI Guide I linked to above (but please note that in at least one case, I had made a mistake and not selected the correct option, as you will see below with Claude in Sections 2 and 3):

  • Claude Opus 4.6 Extended Thinking
  • ChatGPT 5.2 Thinking
  • Gemini 3.0 Pro Deep Research

In addition, in section 3 of this long blogpost, I will very briefly compare and contrast the results I received when I first ran this text prompt through all three GenAI tools on August 7th, 2025, with what I received when I ran them again on Feb. 18th, 2026.

All comparison charts in the February 2026 results in sections 2 and 3 will include some quick stats in a small table under each generative AI tool discussed, namely:

  • the number of characteristics common to all metaverse platforms (and their names); and
  • the number of social VR platforms in the comparison chart (and their names).

Section 4, the final section, contains my overall thoughts after spending a day working with these tools, and a ranking of how well I think these GenAI tools accomplished the given task.


SECTION 2: Comparing Searches Done Feb. 18th, 2026

Feb. 18th, 2026: Claude Opus 4.6 (and Cowork)

First up is Claude. I did this prompt two ways: once via the chatbot interface on the Claude website, and a second time using the Claude app and the new Cowork agent feature. (I was prompted to download and install the Claude app on my Mac, and authenticate using my email address.) First, the chatbot version:

This first report I got back compared eight metaverse characteristics between eight platforms:

8 Metaverse Characteristics8 Social VR Platforms
Persistent Virtual Environments
Real-Time Interactivity
User Identity/Avatars
Social Presence & Co-Experience
User-Generated Content
Virtual Economy
Cross-Platform Accessibility
Interoperability
VRChat
Rec Room
Meta Horizon Worlds
Resonite
Second Life
Spatial
ChilloutVR
NeosVR

Well, right off the bat, I see some problems. First, Second Life is not social VR. Second, it included both Resonite and NeosVR (although Claude told me, “I included both since NeosVR still has historical relevance, but noted it as legacy since the core team transitioned to Resonite”). However, that isn’t a good enough reason to include it in the table.

Then, I turned to the Claude app (which was suggested to me when I did the first text prompt above, so I downloaded and installed it on my MacBook Pro). Then I selected the Cowork (agent) tab along the top three tabs as suggested by Ethan, and I entered the exact same text ptompt:

After beavering away for a few minutes, it gave me the following result:

And when I click on the Open in Firefox button, I get this neatly formatted table (I’m not crazy about the chosen colour scheme, but that’s a minor quibble). It looks good at first:

However, the output, which might look impressive at first, is only as good as the quality of the sources used in its research. If the good information is locked behind a paywall (and therefore, not able to be scraped to add to its knowledge base), then the GenAI tool will use freely-available sources on the web, which can vary quite a bit in quality! There is an acronym in computer science called GIGO: Garbage In, Garbage Out, and I am reminded of this when I decide to take a closer, more critical look at the six sources listed.

All of them were non-academic sources, mostly generic market overviews from websites that I had never heard of before. The six sources included my own list of metaverse platforms on this blog (which is just a list, and doesn’t give any details about the platforms). While I’m flattered they included me, I expected something…more. And I absolutely hated that they mentioned cryptocurrencies, blockchain, DAOs, and NFTs, and included Somnium Space and Decentraland in the resulting table. While Somnium Space is social VR, Decentraland in absolutely not, and I have made my opinions on blockchain-based metaverse platforms very clear in the past on this blog.

8 Metaverse Characteristics6 Social VR Platforms
Persistence
Immersion & Presence
User-Generated Content
Built-In Economy
Social Interaction
Interoperability
Digital Ownership
Decentralized Governance
VRChat
Meta Horizon Worlds
Rec Room
Engage VR
Decentraland
Somnium Space

In fact, I was so dissatisfied with this report that I went back into the Claude Cowork app, and added a qualifier, and made sure that I had turned on Extended Thinking! (I’m almost positive I did that the first time around, but maybe I forgot, and unfortunately, once you’ve done your prompt, the results don’t tell you what modes you used in asking the original question.)

Only to get pretty much the same result: a pretty table with only six websites listed as sources! So much for being more specific and asking for Extended Thinking.

10 Metaverse Characteristics6 Social VR Platforms
Persistence
Immersive 3D Environments
User Identity & Avatars
Real-Time Social Interaction
User-Generated Content
Economy & Monetization
Cross-Platform Access
Scalability & Concurrency
Safety & Moderation
Interoperability
VRChat
Rec Room
Meta Horizon Worlds
Resonite
ChilloutVR
Engage VR

While better thatn the previous round, I am actually disappointed in the results I received from Claude Cowork. But read on; in section 3, I have an update on what I think went wrong here!

Feb. 18th, 2026: ChatGPT 5.2 Thinking

Next, I turned to OpenAI’s ChatGPT, using the ChatGPT 5.2 Thinking mode suggested by Ethan:

And I got back the following table. comparing six social VR platforms on ten metaverse characteristics:

While the resulting table might not be as pretty as the one produced by Claude Opus 4.6 Cowork, I appreciate that there are actual citations which you can hover over and click through to actually see the source material behind the comparison chart entries (and not just a list of websites checked, tacked on to the end). Also, ChatGPT seems to have checked a lot more sources than Claude, and made some sort of attempt to find authoritative sources (often, from the metaverse product’s own online documentation, as shown in this example).

10 Metaverse Characteristics6 Social VR Platforms
Shared Multi-User Spaces
Avatars/Embodied Identity
Real-Time Voice/”Hangout” Core Loop
Persistence (Account, Inventory)
User-Generated Worlds
In-World Creation Tools
Scripting
Economy & Monetization
Cross-Platform Access
Safety Governance
VRChat
Rec Room
Meta Horizon Worlds
Bigscreen Beta
Spatial
Resonite

Overall, I think that ChatGPT 5.2 Thinking gave me a better answer than Claude…but as we will see later on, it doesn’t compare to the best results I got from my day of testing and retesting. Let’s move on to the third of Ethan Mollick’s recommended, general-purpose GenAI tools, Google’s Gemini:

Feb. 18th, 2026: Gemini 3 Pro (first without, and then with, Deep Research)

The first go-round, I selected Gemini 3 Pro mode, as Ethan suggested:

And I got a resulting table comparing three social VR platforms across seven characteristics:

7 Metaverse Characteristics3 Social VR Platforms
Core Philosophy
Visual Style
Creation Tools
Hardware Access
Target Audience
Economy
“Metaverse” Strength (?!)
VRChat
Rec Rooom
Meta Horizon Worlds

I was so unhappy with this first Gemini result that I redid the prompt, this time making sure that I turned on the Deep Thinking mode, just to see if I would get better results, or even some actual citations to sources used:

Wow, what a difference!!

This time around, the task took a lot longer than either Claude or ChatGPT, and it included what appears to be extremely detailed feedback on what was happening behind the scenes (this seems to be turned on by default, and I’m not certain if this mode could have been enabled on Claude or ChatGPT):

And the report I got back was worth the longer wait:

And, at the end, not one but three comparison charts!

Here’s the quick stats, from all three tables in the final report (and notice how technical many of these “metaverse characteristics” are, compared to the other results!):

12 Metaverse Characteristics5 Social VR Platforms
Engine Core
Scripting Language
Persistence Type
Asset Pipeline
Audio Engine
Economic Model
Currency
Identity System
Tracking Support
Instance Cap
Network Model
Culling Tech
VRChat
Rec Room
Roblox
Meta Horizon Worlds
Resonite (only mentioned in one table)

SECTION 3: Comparing August 2025 Prompt Results with the February 2026 Ones

I also wanted to compare the results I when I did the testing last year (August 7th, 2025) with the results I got today (Feb. 18th, 2026) with all three GenAI tools. This was very enlightening.

Then Versus Now: Claude

You will understand why I was so disappointed with today’s results, when you see what the results were when I did the same prompt last year (dated August 7th, 2025):

The report I got back was extremely detailed, with actual citations to sources! I still don’t understand why I got such dramatically different—and worse—results. The difference is so astounding to me that I began to wonder if I had done something wrong this time around.

It was then that I realized that I had literally forgotten to turn on Research mode in the left-hand drop-down menu (previously, I had only had Web Search mode turned on):

So I went to check the Claude app, to see if there was that option available, and, of course, it was there—but under the Chat tab, not the Cowork tab!! So perhaps Cowork still has some user interface bugs to work out. Perhaps sending everything to an agent isn’t the better option; certainly, not in this case!!

Once I had selected both Research and Web Search from the left drop-down menu, and Opus 4.6 Extended from the right drop-down menu, I hit send and waited…until I got a message that I had used up all my credits on my $20-a-month plan!!!

AAAAAAAAAAAAAARGH!!!!

By this point, I was so frustrated with Claude that I simply exited the app. I had had enough frustration for one day.

The next morning, February 19th, 2026, after my daily credits reset at 6:00 pm, I once again tried my prompt with Claude Opus 4.6 Extended Thinking, with both Research and Web Search turned on (using the Gemini app I had installed on my Mac, as opposed to the web version; they appear to be identical in terms of features).

Right off the bat, I got a better response (and Claude even remembered that I was going to working on an OER about the metaverse!):

Again, similar to Google Gemini, I had a bit of wait while Claude did its thing. I actually preferred that Gemini actually gave better descriptions of what it was doing while it was going about its task, as opposed to…well, no updates from Claude other than me sitting and staring at an animated cursor!

Ten minutes later, I got the detailed report I wanted in the first place, and which Claude Cowork stubbornly refused to give me:

The response back included a concise summary taken from the sources examined:

The final report included citations to the academic literature (which I could hover over and click on to go to the source, see the red arrow below), and it cited experts in the field such as Matthew Ball and Tim Sweeney. It’s pretty much all I wanted, and it compares quite favourably to the similarly detailed report from Google Gemini, in the previous section. I am happy.

And this was the only report which had a listing of metaverse characteristics, separate from the ones used in the social VR platforms comparison chart:

Here’s the quick stats from the comparison chart. As you can see, there are some problems here, with the inclusion of platforms which are clearly not social VR (e.g. Second Life) and platforms that no longer exist (Altspace shut down on March 10th, 2023). These sort of mistakes make we wonder about the accuracy and currency of the report overall.

9 Metaverse Characteristics9 Social VR Platforms
Persistence
Synchronous Real-Time
Massive Scale/Concurrency
Cross-Platform Access
Virtual Economy
User-Generated Content
Interoperability
Avatar/Identity Systems
Immersive 3D/Spatial Computing
Open Standards/Decentralization
Spanning Physical-Digital
Ethical Goivernance/Accessibility
VRChat
Horizon Worlds (note: old name used)
Rec Room
Resonite
ChilloutVR
AlspaceVR (was shut down)
Second Life (not social VR!)
Roblox
Fortnite (not social VR!)

Then Versus Now: ChatGPT

An interesting difference between the August 2025 report from ChatGPT and today’s report is this: in last year’s report, for whatever reason, the tool asked me a follow-up question to clarify what was wanted (I did use the Deep Research feature in the 2025 report, as well):

Based on that clarification prompted by ChatGPT, I actually think I preferred the 2025 report format over this new one. So why didn’t ChatGPT 5.2 Thinking ask me any follow-up questions this time around? And that’s part of the frustration with tese tools; the way that they operate is still very much a black box, where you don’t understand how the tool is processing what you ask of it.

Then Versus Now: Gemini

The last comparison is between the Google Gemini report I produced on August 7th, 2025, and today’s report. One thing I noticed about the Aug. 7th report is how hard it tried to shoehorn in an overarching narrative into the final result, in a way that seemed a bit hamfisted, frankly. But the result was still a very detailed report with an extensive list of citations, comparable to today’s report. I prefer today’s version.


SECTION 4: Executive Summary and Ranking

This is going to be concise, I promise! Five points.

First, while we might be entering what Ethan Mollick calls “the agentic era,” my experience today shows that simply handing something off to an agent, as opposed to the back-and-forth conversation with a chatbot interface, does not always give the best result. In particular, Claude Cowork gave me terrible results, and eventually, I ran out of daily use credits to actually run the report I wanted in the first place.

Second, the user interface for these GenAI tools is awful and NON-intuitive. Hiding critical options like Deep Research under drop-down menus, and not making it clear what options have been selected when you do a text prompt, is a major problem. All three companies need to hire some good user interface/user experience staff. If I, with decades of computer experience and a goddamn computer science degree, can’t figure this shit out, God help the average non-technical user—and isn’t that what the point of generative AI is supposed to be, to make it easier for the user to do things??

Third, when these tools work, they are astoundingly good (the Gemini 3.0 Pro report with Deep Research turned on, and the Claude Opus 4.6 report with Research, Web Search, and Extended Thinking turned on). But when they don’t, they can still fail spectacularly (Claude Cowork). So you still have to be the human in the loop here, to figure out when you get a good result versus a bad one. What is frustrating is that all these GenAI tools operate in a black box, with only Gemini making some attempt at explaining what it was doing, as it was doing it.

Fourth, as Ethan himself said in his latest AI Guide:

The top models are remarkably close in overall capability and are generally “smarter” and make fewer errors than ever. But, if you want to use an advanced AI seriously, you’ll need to pay at least $20 a month (though some areas of the world have alternate plans that charge less). Those $20 get you two things: a choice of which model to use and the ability to use the more advanced frontier models and apps. I wish I could tell you the free models currently available are as good as the paid models, but they are not.

In other words, you get what you pay for. And sometimes, even the $20-a-month level isn’t enough, as seen with my experience on Feb. 18th with Claude (and yes, using the cutting-edge features does eat into your usage limits pretty quickly, as I learned to my chagrin).

Finally, I have found that the one of the best ways to see where the strengths and weaknesses of these GenAI tools is to enter the exact same text prompt into each of them, and then compare and contrast the results you get back. However, that approach is gonna cost you at least US$60 a month, so it might not be worth it to you. (And will I be doing this forever? No; at some point, I will just pick one or perhaps two tools and cancel my subscriptions to the rest of them.)

So, in this current round of testing, I would rank the results as follows (separating the results from Claude into the chatbot-generated report and the Cowork report):

  1. Google Gemini 3.0 Pro (with Deep Research turned on) provided me with a very detailed report with citations, as well as giving me a detailed play-by-play on how it was answering my query, which I really appreciated.
  2. Claude Opus 4.6 report (with Research, Web Search, and Extended Thinking turned on) also gave me a detailed report with citations, but several errors in the comparison chart made me question the overall quality and currency of the report. I also really hated how I had to futz around to get the results I really wanted!
  3. ChatGPT 5.2 Thinking is in a clear third place, in my opinion. Not bad, but not as detailed a result as Gemini and Claude provided.
  4. Claude Opus 4.6 Cowork, with perhaps the prettiest output but easly the least substantial result, using lower-quality sources of information, clearly failed at this task. For those reasons, I ranked it in last place. Ethan’s “Agentic Era” might be true for some applications, but certainly not this one!

I have found these little excursions into generative AI to be quite enlightening, and they have definitely given me some new ideas of topics to explore when I begin my research and study leave to write an OER about the metaverse. Hopefully, you found it enlightening, too. Please go subscribe to Ethan Mollick’s free Substack newsletter; he tends to update his AI Guide recommendations fairly regularly, and it’s really the best way too stay on top of a rapidly changing and evolving field!

UPDATED TO VERSION 1.3! Second Life Steals, Deals, and Freebies: A New Comparison Chart of Seven Options for Free or Inexpensive Female Mesh Bodies (Including Senra Jamie)

Now that Linden Lab has launched the beta version of its Senra mesh starter avatars, I decided to take a stab at creating a comparison chart, comparing and contrasting six options for free or inexpensive (L$250 or less) female mesh bodies. (I will probably follow up with a similar chart for free/inexpensive male mesh bodies, female mesh heads, and male mesh heads.)

The six seven mesh bodies I have chosen for this chart are:

  • Senra Jamie, by Linden Lab (UPDATE March 16th, 2024: These are now out of beta test, and finalized to version 1.0)
  • Erika Zero X, by Kalhene
  • Atenea, by LucyBody
  • Classic Meshbody (often referred to as TMP)
  • eBody Classic (free version)
  • eBody Curvy (free version)
  • UPDATE Aug. 4th, 2023: After some hemming and hawing, I have decided to include the open-source Ruth 2.0 mesh bodies in this spreadsheet. You can find a list of vendors for Ruth 2.0-based mesh bodies here (scroll down to the Ruth2 section). While clothing specifically designed for Ruth 2.0 bodies is limited, with a good set of BoM/system alphas, some Maitreya Lara clothing and standard-size clothing does fit.

UPDATE March 15th, 2024: You used to have to pay to join the Erika Mesh Body group to pick up the free group gift of the Erika Zero X mesh body, but you can now join the group for free! I have updated my comparison chart to version 1.3 with this updated information. This is a lovely female mesh body which responds very well to the body sliders, allowing you a wide variety of body shapes, from thin and slim to “thicc” and curvy!

For each mesh body, I look at the following:

  • Price
  • Bakes on Mesh support
  • Bento support
  • Feet options and compatibility
  • Mesh clothing compatibility (please note that all Bakes on Mesh bodies support BoM/system layer clothing; here are some places where you can find those)
  • Mesh clothing availability (obviously a subjective estimate!)

You can view (but not edit) version 1.3 of my comparison chart here on Google Drive. I am open to suggestions for improving this chart, and I expect to keep it (somewhat) updated as the situation evolves over time. If you have any corrections, edits, or suggestions, please leave a comment, thanks!

Here’s a snapshot of version 1.3 of the comparison chart, which you can view and download in full size over on Flickr if you, like me, find the fine print a little too small:

Comparison Chart of Free and Inexpensive Female Mesh Bodies 16 March 2024

Please note that I have deliberately excluded some mesh bodies, for example, the free Altamura bodies you can pick up at various locations (because you cannot change the skin, and you cannot use Bakes on Mesh with them). I have also left out those bodies which have poor or even non-existent third-party designer support. An example of this would be the Ultra Vixen mesh body, which is now only free to avatars under 30 days old and—as far as I am aware—only has clothing that fits it, which is made and sold by the body’s creator.

Looking forward to hearing your comments and suggestions!

Comparing and Contrasting Three Artificial Intelligence Text-to-Art Tools: Stable Diffusion, Midjourney, and DALL-E 2 (Plus a Tantalizing Preview of AI Text-to-Video Editing!)

HOUSEKEEPING NOTE: Yes, I know, I know—I’m off on yet another tangent on this blog! Please know that I will continue to post “news and views on social VR, virtual worlds, and the metaverse” (as the tagline of the RyanSchultz.com blog states) in the coming months! However, over the next few weeks, I will be focusing a bit on the exciting new world of AI-generated art. Patience! 😉

Artificial Intelligence (AI) tools which can create art from a natural-language text prompt are evolving at such a fast pace that it is making me a bit dizzy. Two years ago, if somebody had told me that you would be able to generate a convincing photograph or a detailed painting from a text description alone, I would have scoffed! Many felt that the realm of the artist or photographer would be among the last holdouts where a human being was necessary to produce good work. And yet, here we are, in mid-2022, with any number of public and private AI initiatives which can be used by both amateurs and professionals to generate stunning art!

In a recent interview by The Register‘s Thomas Claburn of David Holz (the former co-founder of augmented reality hardware firm Magic Leap, who founded Midjourney), there’s a brief explanation of how this burst of research and development activity got started:

The ability to create high-quality images from AI models using text input became a popular activity last year following the release of OpenAI’s CLIP (Contrastive Language–Image Pre-training), which was designed to evaluate how well generated images align with text descriptions. After its release, artist Ryan Murdock…found the process could be reversed – by providing text input, you could get image output with the help of other AI models.

After that, the generative art community embarked on a period of feverish exploration, publishing Python code to create images using a variety of models and techniques.

“Sometime last year, we saw that there were certain areas of AI that were progressing in really interesting ways,” Holz explained in an interview with The Register. “One of them was AI’s ability to understand language.”

Holz pointed to developments like transformers, a deep learning model that informs CLIP, and diffusion models, an alternative to GANs [Holz pointed to developments like transformers, a deep learning model that informs CLIP, and diffusion models, an alternative to GANs [models using Generative Adversarial Networks]. “The one that really struck my eye personally was the CLIP-guided diffusion,” he said, developed by Katherine Crawson…

If you need a (relatively) easy-to-understand explainer on how this new diffusion model works, well then, YouTube comes to your rescue with this video with 4 explanations at various levels of difficulty!


Before we get started, a few updates since my last blogpost on A.I.-generated art: After using up my free Midjourney credits, I decided to purchase a US$10-a-month subscription to continue to play around with it. This is enough credit to generate approximately 200 images per month. Also, as a thank you for being among the early beta testers of DALL-E 2, the AI art-generation tool by OpenAI, they have awarded me 100 free credits to use. You can buy additional credits in 115-generation increments for US$15, but given the hit-or-miss nature of the results returned, this means that DALL-E 2 is among the most expensive of the artificial intelligence art generators. It will be interesting to see if and how OpenAI will adjust their pricing as the newer competitors start to nip at their heels in this race!

And I can hardly believe my good fortune, because I have been accepted into the relatively small beta test group for a third AI text-to-art generation program! This new one is called Stable Diffusion, by Stability AI. Please note that if you were to try to get into the beta now, it’s probably too late; they have already announced that they have all the testers they need. I submitted my name 2-3 weeks ago, when I first heard about the project. Stable Diffusion is still available for researcher use, however.

Like Midjourney, Stable Diffusion uses a special Discord server with commands (instead of Midjourney’s /imagine, you use the prompt !dream, followed by a text description of what you want to see, plus you can add optional parameters to set the aspect ratio, the number of images returned, etc.). However, the Stable Diffusion team has already announced that they plan to move from Discord to a web-based interface like DALL-E 2 (we will be beta-testing that, too). Here’s a brief video glimpse of what the web interface could look like:


Given that I am among the relatively few people who currently have access to all three of the top publicly-available AI art-generation tools, I thought it would be interesting to create a chart comparing and contrasting all three programs. Please note that I am neither an artist nor an expert in artificial intelligence, just a novice user of all three tools! Almost all of the information in this chart has been gleaned from the websites of the projects, and online news reports, as well as the active subreddit communities for all three programs on Reddit, where users post pictures and ask questions. Also, all three tools are constantly being updated, so this chart might go very quickly out-of-date (although I will make an attempt to update it).

Name of ToolDALL-E 2MidjourneyStable Diffusion
CompanyOpenAIMidjourneyStability AI
AI Model UsedDiffusionDiffusionDiffusion
# Images Used
to Train the AI
400 millon“tens of millions”2 billion
User InterfacewebsiteDiscordDiscord (moving to website)
Cost to Usecredit system (115 for US$15)subscription (US$10-30 per month)currently free (beta)
Uses Text Promptsyesyesyes
Can Add Optional Argumentsnoyesyes
Non-Square Images?noyesyes
In-tool Editing?yesnono
Uncropping?yesnono
Generate Variations?yesyesyes (using seeds)
A comparison chart of three AI text-to-art tools: DALL-E 2, Midjourney, and Stable DIffusion

I have already shared a few images from my previous testing of DALL-E 2 and Midjourney here, here, and here, so I am not going to repost those images, but I wanted to share a couple of the first images I was able to create using Stable Diffusion (SD). To make these, I used the text prompt “a thatched cottage with lit windows by a lake in a lush green forest golden hour peaceful calm serene very highly detailed painting by thomas kinkade and albrecht bierstadt”:

I must admit that I am quite impressed by these pictures! I had asked SD for images with a height of 512 pixels and a width of 1024 pixels, but to my surprise, the second image was a wider one presented neatly in a white frame, which I cropped using my trusty SnagIt image editor! Also, it was not until after I submitted my prompt that I realized that the second artist’s name is actually ALBERT Bierstadt, not Albrecht! It doesn’t appear as if my typo made a big difference in the final output; perhaps for well-known artists, the last name alone is enough to indicate a desired art style?

Here are a few more samples of the kind of art which Stable Diffusion can create, taken from the pod-submissions thread on the SD Discord server:

Text prompt: “a beautiful landscape photography of Ciucas mountains mountains a dead intricate tree in the foreground sunset dramatic lighting by Marc Adamus”
Text prompt: “incredible wide screenshot ultrawide simple watercolor rough paper texture katsuhiro otomo ghost in the shell movie scene backlit distant shot”
Text prompt: “an award winning wallpaper of a beautiful grassy sunset clouds in the sky green field DSLR photography clear image”
Text prompt: “beautiful angel brown skin asymmetrical face ethereal volumetric light sharp focus”
Painting of people swimming (no text prompt shared)

You can see many more examples over at the r/StableDiffusion subreddit. Enjoy!

If you are curious about Stable Diffusion and want to learn more, there is a 1-1/2 hour podcast interview with Emad Mostaque, the founder of Stability AI (highly recommended!). You can also visit the Stability AI website, or follow them on social media: Twitter or LinkedIn.


I also wanted to submit the same text prompt to each of DALL-E 2, Midjourney, and Stable Diffusion, to see how the AI models in each would respond. Under each prompt you will see three square images: the first from DALL-E 2, the second from Midjourney, and the third from Stable Diffusion. (Click on each thumbnail image to see it in its full size on-screen.)

Text prompt: “the crowds at the Black Friday sales at Walmart, a masterpiece painting by Rembrandt van Rijn”

Note that none of the AI models are very good at getting the facial details correct for large crowds of people (all work better with just one face in the picture, like a portrait, although sometimes they struggle with matching eyes or hands). I would say that Midjourney is the clear winner here, although a longer, much more detailed prompt in DALL-E 2 or Stable Diffusion might have created an excellent picture.

Text prompt: “stunning breathtaking photo of a wood nymph with green hair and elf ears in a hazy forest at dusk. dark, moody, eerie lighting, brilliant use of glowing light and shadow. sigma 8.5mm f/1.4”

When I tired to generate a 1024-by-1024 image in Stable Diffusion, it kept giving me more than one wood nymph, even when I added words like “single” or “alone”, which is a known bug in the current early state of the program. I finally gave up and used a 512×512 image. The clear winner here is DALL-E 2, which has a truly impressive ability to mimic various camera styles and settings!

Text prompt: “a very highly detailed portrait of an African samurai by Tim Okamura”

In this case, the clear winner is Stable Diffusion with its incredible detail, even though, once again, I could not generate a 1024×1024 image because it kept giving me multiple heads! The DALL-E 2 image is a too stylized for my taste, and the Midjourney image, while nice, has eyes that don’t match (a common problem with all three tools).

And, if you enjoy this kind of thing, here’s a 15-minute YouTube video with 21 more head-to-head comparisons between Stable Diffusion, DALL-E 2, and Midjourney:


As I have said, all of this is happening so quickly that it is making my head spin! If anything, the research and development of these tools is only going to accelerate over time. And we are going to see this technology applied to more than still images! Witness a video shared on Twitter by Patrick Esser, an AI research scientist at Runway, where the entire scene around a tennis player is changed simply by editing a text prompt, in real time:


I expect I will be posting more later about these and other new AI art generation tools as they arise; stay tuned for updates!

A.I.-Generated Art: Comparing and Contrasting DALL-E 2 and Midjourney as Both Tools Move to an Open Beta

UPDATE Aug. 12th, 2022: I have just joined the beta test of Stable Diffusion, another AI art-generation program! For more information, please read Comparing and Contrasting Three Artificial Intelligence Text-to-Art Tools: Stable Diffusion, Midjourney, and DALL-E 2 (Plus a Tantalizing Preview of AI Text-to-Video Editing!)

You might remember that I was one of the lucky few who received an invitation to be part of the closed beta test (or “research preview”, as they called it) of DALL-E 2, a new artificial intelligence tool from a company called OpenAI, which can create art from a natural-language text prompt. (I blogged about it, sharing some of the images I created, here and here.)

Here are a few more pictures I generated using DALL-E 2 since then (along with the prompt text in the captions):

DALL-E 2 prompt: “feeling despair over a uncertain future digital art”
DALL-E 2 prompt: “feeling anxiety over an uncertain future digital art”
DALL-E 2 prompt: “feeling anxiety over a precarious future” (sensing a theme here?)
DALL-E 2 prompt” “award-winning detailed vibrant bright colorful knife painting by Françoise Nielly” (Note that this used an inpainting technique; I expanded the canvas borders and asked DALL-E 2 to fill them in to match the Nielly knife painting of the man’s face in the middle)

Meanwhile, other DALL-E 2 users have generated much better results than I could, by skillful use of the text prompts. Here are just a few examples from the r/dalle2 subReddit community of AI-generated images which impressed and sometimes even stunned me, with a direct link to the posts in the caption underneath each picture:

DALL-E 2 prompt: “an image of the Cosmic Mind, digital art”
DALL-E 2 pompt: “cyborg clown, CGSociety award winning render”
DALL-E 2 prompt: “a young girl stares directly at the camera, her blue hijab framing her face. The background is a blur of colours, possibly a market stall. The photo is taken from a low angle, making the girl appear vulnerable and child-like. Kodak Portra 400”
DALL-E 2 prompt: “a close-up photograph of a man with brown hair, ice-blue eyes, red and brown stubble Balbo beard, his face is narrow, with defined cheekbones, he has a scar on the left side of his lips, running down from his top to the bottom lip, he wears a dark-blue hoodie, the background is a blurred out city-scape”

As you can see by the last two images, you can get very detailed and technical in your text prompts, even including the model of camera used! (However, also note that in the fourth picture, DALL-E 2 ignored some specific details in the prompt.)

Yesterday, OpenAI sent me an email to annouce that DALL-E 2 was moving into open beta:

Our goal is to invite 1 million people over the coming weeks. Here’s relevant info about the beta:

Every DALL·E user will receive 50 free credits during their first month of use, and 15 free credits every subsequent month. You can buy additional credits in 115-generation increments for $15.

You’ll continue to use one credit for one DALL·E prompt generation — returning four images — or an edit or variation prompt, which returns three images.

We welcome feedback, and plan to explore other pricing options that will align with users’ creative processes as we learn more.

As thanks for your support during the research preview we’ve added an additional 100 credits to your account.

Before DALL-E 2 announced their new credits system, I had spent most of one day’s free prompts during the research preview to try and generate some repeating, seamless textures to apply to full-permissions mesh clothing I had purchased from the Second Life Marketplace. Most of my attempts were failures, pretty designs but not 100% seamless. However, I did manage to create a couple of floral patterns that worked:

So, instead of purchasing texture packs from without and outside of Second Life, I could, theoretically, generate unique textile patterns, apply them to mesh garments, and sell them, because according to the DALL-E 2 beta announcement I received:

Starting today, you get full rights to commercialize the images you create with DALL·E, so long as you follow our content policy and terms. These rights include rights to reprint, sell, and merchandise the images.

You get these rights regardless of whether you used a free or paid credit to generate images, and this includes images you’ve created before today during the research preview.

Will I? Probably not, because it took me somewhere between 20 and 30 text prompts to generate only two useful seamless patterns, so it’s just not cost effective. However, once AI art tools like DALL-E 2 learns how to generate seamless textures, it’s probably going to have some sort of impact on the texture industry, both within and outside of Second Life! (I can certainly see some enterprising soul set up a store and sell AI-generated art in a virtual world; SL is already full of galleries with human-generated art.)


Another cutting-edge AI art-generation program, called Midjourney (WARNING: ASCII art website!), has also announced an open beta. I had signed up to join the waiting list for an invitation several weeks ago, and when I checked my email, lo and behold, there it was!

Hi everyone,

We’re excited to have you as an early tester in the Midjourney Beta!

To expand the community sustainably, we’re giving everyone a limited trial (around 25 queries with the system), and then several options to buy a full membership.

Full memberships include; unlimited generations (or limited w a cheap tier), generous commercial terms and beta invites to give to friends.

Although both DALL-E 2 and Midjourney use human text prompts to generate art, they operate differently. While DALL-E 2 uses a website, Midjourney uses a special Discord server, where you enter your prompt as a special command, generating four rough thumbnail images, which you can then choose to upscale to a full-size image, or use as the basis for variations.

I took some screen captures of the process, so you can see how it works. I typed in “/imagine a magnificent sailing ship on a stormy sea”, and got this back:

The U buttons will upscale one of the four thumbnails, adding more details, while the V buttons generate variations, using one of the four thumbnails as a starting point. I choose thumbnail four and generated four variations of that picture:

Then, I went back and picked one of my original four images to upscale. You can actually watch as Midjourney slowly adds details to your image, it’s fascinating!

I then clicked on the Upscale to Max button, to receive the following image:

My first attempt at generating an image using Midjourney

Now, I am not exactly satisfied with this first attempt (that sailing ship looks rather spidery to me), but as with DALL-E 2, you get much better results with more specific, detailed text prompts. Here are a few examples I took from the Midjourney subReddit community (with links back to the posts in the captions):

Midjorney prompt: “cyberpunk soldier piloting a warship into battle, the atmosphere is like war, fog, artstation, photorealistic”
Midjourney prompt: “Dress made with flowers” (click to see a second one on Reddit)
Midjourney prompt: “a tiny stream of water flows through the forest floor, octane render, light reflection, extreme closeup, highly detailed, 4K

So, as you can see, you can get some pretty spectacular results, with incredible levels of detail! And unlike DALL-E 2, you can set the aspect ratio of your pictures (as was done in the fourth image generated). You do this with a special “–ar” command in your text prompt to Midjourney, e.g. “–ar 16:9” (here’s the online documentation explaining the various commands you can use).

And one area in which Midjourney appears to excel is horror:

Midjourney prompt: “a pained, tormented mind visualized as a spiraling path into the void”
Midjourney prompt: “a beautiful painting of Escape from tarkov in machinarium style, insanely detailed and intricate, golden ratio, hypermaximalist, elegant, ornate, luxury, elite, horror, creepy, ominous, haunting, matte painting, cinematic, cgsociety, James jean, Brian froud, ross tran”

You can see many more examples of depictions of horror in the postings to the Midjourney SubReddit; some are much creepier than these!


So, in comparing the two tools, I think that Midjourney offers more parameters to users (e.g. setting an aspect ratio), which DALL-E currently lacks. Midjourney also seems to produce much more detailed images than DALL-E 2 does, whereas DALL-E 2 is often astoundingly good at a much wider variety of tasks. For example, how about some angry bison logos for your football team?

I think these images are all very good! (Note that DALL-E 2 still struggles with text! Midjourney does too, but it gets the text correct more often than DALL-E 2 does at present. But note that might change over time as both systems evolve.)


So, the good news is that both DALL-E 2 and Midjourney are now in open beta, which means that more people (artists and non-artists alike) will get an opportunity to try them out. The bad news is that both still have long waiting lists, and with the move to beta, both DALL-E 2 and Midjourney have put limits in place as to how many free images you can generate.

Midjourney gives you a very limited trial period (about 25 prompts), and then urges you to pay for a subscription, with two options:

Basic membership gives you around 200 images per month for US$10 monthly; standard membership gives you unlimited use of Midjourney for US$30 a month.

For now, OpenAI has decided to set DALL-E 2’s pricing based on a credit system (similar to their GPT-3 AI text-generation tool), as described in the first quote in this blogpost. There’s no option for unlimited use of DALL-E 2 at any price, just options for buying credits in different amounts (and there are no volume discounts for purchasing larger amounts of credits at one time, either). The most you can by at once is 5,750 credits, which is US$750. So, yes, it can get quite expensive! (As far as I am aware, your unused credits carry over from one month to the next.)

There’s quite a bit of discussion about OpenAI’s DALL-E 2 pricing model in this thread in the r/dalle2 subReddit; many people are very unhappy with it, particularly since it can take a lot of trial and error with the DALL-E 2 text prompts to generate a desired result. One person said:

In my experience, using Dall-E 2 to generate concept arts for our next project, it takes me between 10 to 20 attempts to get something close to what I want (and I never got exactly what I was asking for)…

Dall-E 2, at this point, is not a professional tool. It’s not viable as one, unless you produce exactly the type of content the AI can produce instantly just the way you want it.

Dall-E 2, at this point, IS A TOY! And that’s OpenAI’s mistake right now. You can’t sell a toy the way you sell a professional service! I’m ready to pay for it because I’m experimenting with it. I’m having fun with it and, when it works, it provides me with images I can also use for professional project. However, I wont EVER spend hundreds of dollars on this just for fun, and I certainly wont pay that amount for it as a tool until it can provide me with better and more consistent results!

OpenAI is going after the WRONG TARGET! OpenAI should be seeling it at a much lower price for everyday people and enthusiasts who want to experiment with it because this is literally the only people who can be 100% satisfied with it at this point and these people wont pay hundreds of dollars per month to keep playing when there are other shiny toys out there, cheaper and more open, existing or about to.

Several commenters said that they will be moving from DALL-E 2 to Midjourney because of their more favourable pricing model, but of course it’s still early days. Also, there are any number of open-source AI art-generation projects in the works, and competition will likely mean more features (and better results!) at less cost over time. One thing is certain: we can anticipate an acceleration in improvement of these tools over time.

The future looks to be both exciting and scary! Exciting in the ability to generate art in a new way, which up until now has been restricted to experienced artists or photographers, and scary in that we can no longer trust our eyes that a photograph is real, or has been generated by artificial intelligence! Currently, both systems have rules in place to prevent the creation of deepfake images, but in future, things could get Black Mirror weird, and the implications to society could be substantial. (Perhaps now you will understand the first three DALL-E 2 text prompts I used, at the top of this blogpost!)

P.S. Fun fact: the founding CEO of Linden Lab (the makers of Second Life), Philip Rosedale, is one of the advisors to Midjourney, according to their website. Philip gets around! 😉

UPDATE July 22nd, 2022: Of course, the images generated by DALL-E 2 and Midjourney can then be used in other AI tools, such as WOMBO and Reface (please click the links to see all the blogposts I have written about these mobile apps).

Late yesterday, a member of the r/dalle2 community posted the following 18-second video, created by generating a photorealistic portrait of a woman using DALL-E 2, then submitting it to a tool similar to WOMBO and Reface called Deep Nostalgia:

What you see here is an AI-generated image, “animated” using another deep learning tool. This is a tantalizing glimpse into the future, where artificial intelligence can not only create still images, but eventually, video!