← Back to model catalog A simple illustration of the Qwen 3.5 Cloud setup -- your computer talks to OpenWebUI, which talks to Qwen 3.5 running in the cloud

Image source: Ollama

Qwen 3.5 · Terminal Glass

A Do-It-All AI That Reads Text and Images -- Without Buying Powerful Hardware

Qwen 3.5 runs through a simple chat window on a modest computer, while the actual AI work happens on Ollama's cloud servers. It can read both text and images, work in 201 languages, and take in entire lengthy documents in a single conversation. You avoid buying expensive computing hardware for a model this large -- in exchange, your questions, any documents you search, and any images you upload travel out to those cloud servers, and you're billed based on how much you use it.

In short: your computer runs the chat window. Qwen 3.5, running on Ollama's cloud, does the heavy lifting.

So, what exactly is Qwen 3.5?

Qwen 3.5 is Alibaba Cloud's AI model family that can understand both text and images at the same time, follow instructions closely, help with coding, use outside tools when needed, and communicate in 201 languages and dialects. Some versions are small enough to run on your own laptop; the cloud versions let you use the largest, most capable versions without downloading huge files or owning a powerful graphics card yourself.

Through the cloud, you can access a standard version as well as an even larger, more capable version -- both of which can handle entire long documents in one conversation and understand images you upload. It's a strong, well-rounded, everyday choice. If your main focus is heavy-duty coding across large codebases, you may still prefer a coding-specialized model instead.

Here's the honest trade-off

Terminal Glass gives you the same friendly chat app (OpenWebUI) and the same locally stored document search tool you'd get with a fully private setup -- but instead of your own computer doing the "thinking," each question gets sent out to Ollama's cloud to be answered by Qwen 3.5. Your own machine just needs to run the chat window and your document search -- a small, inexpensive computer is usually enough. You don't need to store or run a massive AI model file on your own hardware.

Here's what that really means: every prompt, every document snippet your search tool pulls up, and any image you upload gets sent out over the internet to be processed. This is not a locked-down, private setup -- it's a different thing entirely from NoCloudGPT Private Instances, where everything stays on machines you own and control. You'll also pay an ongoing, usage-based fee that grows as your team, your documents, and your chat volume grow. If privacy matters more to you than convenience, choose a NoCloudGPT Air Gapped setup instead.

"Do I need to buy a fancy graphics card?"

This is really the question most people are trying to figure out, even if they don't phrase it that way.

If you try to run these AI models without a proper graphics card (the same kind used for video games, but scaled up), your computer will fall back to using its regular processor. That technically works for the occasional light question, but it becomes painfully slow for regular use. A lot of people learn this the hard way after renting a cloud computer that turns out to be either too slow or surprisingly expensive.

If you'd rather not deal with picking out graphics cards, comparing specs, or troubleshooting driver issues at all, Terminal Glass Cloud is the easier path -- you get a capable AI through a simple, familiar chat window, and someone else worries about the hardware.

On the other hand, if you want everything running on equipment you personally control, you can rent a graphics-card-equipped server (through AWS, Digital Ocean, or your own machines) and install NoCloudGPT on it yourself. In that setup, you're the one responsible for picking the right graphics card and keeping it running smoothly.

And if your organization needs everything to stay completely inside your own walls from the very first day, then a NoCloudGPT Air Gapped setup on your own hardware is the way to go.

There's no single "correct" choice here. It really comes down to your budget, how sensitive your data is, how fast you need answers, and whether you want to own your equipment long-term.

Is this actually for you?

👤 If you're using this for yourself

A practical fit if you want help drafting writing, summarizing long documents, translating between languages, or analyzing a screenshot -- but your own computer isn't powerful enough to run a large AI model on its own. The cloud version can take in entire lengthy reports in one go. Just remember that anything you type or upload travels out to the cloud, so keep personal journals, financial details, and other private material on a fully private setup instead.

🏢 If you're using this for your business or team

Small teams use this for internal drafting, meeting summaries, multilingual customer communication, and searching through company documents via a shared chat tool. It's especially good at following instructions precisely and translating accurately -- useful for teams working across countries. It is not appropriate for medical records, attorney-client material, sensitive financial data, or anything that legally must never leave your network -- those need a fully private setup instead.

⚙️ If you're technical or building software

Useful for testing out AI-powered features, generating starter code, getting code explained to you, and experimenting with AI "tool use" before you commit to buying or renting serious hardware. Because it uses the same standard tools underneath, a project built this way can later be moved to a fully private, self-hosted setup with minimal rework. Just don't send proprietary source code you can't risk exposing to outside servers.

🔬 If you're doing research or working with data

Handy for exploring research literature, sorting and labeling text, brainstorming hypotheses, and cleaning up messy public datasets when you want strong performance without going through a hardware-buying process. Keep in mind a cloud model can change or update behind the scenes, so it won't give you the exact same, repeatable results every time. Confidential data or anything tied to a publication should stay on a fully private setup with a fixed, unchanging version of the model.

How this actually works, step by step

You'll interact with everything through OpenWebUI -- a clean, simple chat window running on your own computer (or you can use the command line, if you prefer). Your chat history, user accounts, and any documents you've indexed for search stay on hardware you control. But the moment you send a question or attach an image, it travels out to Ollama's cloud servers to be processed by Qwen 3.5, and the answer streams back to your screen.

To use the cloud version, you'll need to sign in with an Ollama account on your computer first. You're billed based on how much you use it -- and the larger version, along with longer conversations, costs more per request than shorter chats on the standard version.

ollama signin
ollama run qwen3.5:cloud

What leaves your computer: your typed questions, any document text your search tool pulls up, and any images you submit for analysis. Speed: usually a bit slower than a model running on your own machine, and depends on your internet connection and how busy the cloud servers are. Cost: based on how much you use it -- worth setting expectations early if several people share one setup or you're running things automatically.

When you should probably choose something else

What we recommend

Version Best for Notes
qwen3.5:cloud Individuals and small teams Handles very long documents and understands images. The default, everyday choice for chatting, drafting, document search, and light experimentation, in any of 201 languages. Get quote →
qwen3.5:397b-cloud Demanding reasoning and multi-step tasks Same long-document and image capabilities, but the largest and most capable version available. Costs more per use -- worth it when the extra quality clearly pays off. Get quote →

Terminal Glass sets up your chat window, your document search tool, and helps you keep an eye on your cloud usage. What's available can change over time -- double-check the current options at ollama.com/library/qwen3.5 before committing.

The legal fine print, in plain terms

Using the cloud versions of Qwen 3.5 requires an Ollama Cloud account, and you're billed based on how much you use it -- usage limits and pricing tiers apply and can change. The downloadable versions of Qwen 3.5 are generally released under the Apache 2.0 license, which allows commercial use as long as you follow its standard terms. Always check the license for the specific version you're using before relying on it for business purposes -- some cloud-only versions from Alibaba are separate products with their own, different terms.

Official listing: ollama.com/library/qwen3.5 · Model files: Qwen on Hugging Face

A quick disclaimer: Terminal Glass is our own way of packaging this setup -- we don't own or run Qwen 3.5 or Ollama Cloud ourselves. Availability, pricing, and behavior are controlled by Alibaba Cloud and Ollama, and can change without notice. This page is meant to help you understand your options -- it isn't legal advice, and you're responsible for making sure your use complies with any relevant license terms or data-protection rules.
Get a quote for Qwen 3.5 Cloud →

We'll help you figure out the right small computer to run locally and the right cloud option for your team.

Image and source credits

Hero image

Shared with the local Qwen 3.5 page. Sourced from the Ollama library listing.

Who makes this model

Qwen 3.5 is made by Alibaba Cloud's Qwen team. Terminal Glass and NoCloudGPT are separate services and aren't affiliated with Alibaba or Ollama.

Want it fully private instead?

See Qwen 3.5 -- NoCloudGPT Private Instances for a version that runs entirely, air-gapped, on hardware you control.