In brief
- Images can be sent to the Gemini API via a URL, inline data, or the Files API.
- Requests using inline image data have a total size limit of 20 MB.
- The Files API is recommended for large files or images that will be reused.
- Clear, specific instructions help define the desired response format and scope.
What will you learn in this guide?
You can send an architectural photo or render to a model and ask it to describe visible elements, classify what it can see, or answer focused questions for design evaluation. This guide walks through sending a single image to a model via the Gemini API and creating prompts to try in architecture, interior design, and visualization workflows.
The Gemini documentation covers multimodal tasks such as image description, classification, and visual question answering. The architectural examples below are suggested prompts based on these tasks; the sources do not guarantee the accuracy or suitability of each specific use case listed here for a particular workflow. A model response is not a substitute for project drawings or technical verification. Think of it as a starting point for organizing information from an image and supporting internal team reviews.
Requirements and preparing your image
The examples use the gemini-3.8-flash model in the Gemini API. Code samples in the source documentation show how to make requests using Python, JavaScript, Java, Go, and REST, but no specific SDK version is given here. In the REST example, the API key is used through the GEMINI_API_KEY variable. Check setup and access details separately for your development environment.
First, choose a single image to analyze: it could be an interior render, a façade image, a photo of a site plan, or material samples. Clearly state what you expect the model to do in your prompt. For example, ask it to list only visible elements, focus on a particular visual feature, or flag uncertain details.
You can send an image to the API in three ways: an accessible image URL, inline (base64-encoded) image data, or the Files API. When using inline data, the total request size—including text prompts and other content—must not exceed 20 MB. The Files API is recommended for large files or when the same file will be used in multiple requests.
Step by step: five prompts for image analysis
1. List what is visible in the render objectively. Send the image to the API alongside text and add the prompt below. Focusing on observable elements rather than interpretation in an initial review can provide a more organized response for later evaluation.
Bu mimari görselini incele. Yalnızca görüntüde seçilebilen öğeleri belirt: mekân türü, ana yüzeyler, mobilyalar, aydınlatma ve bitkiler. Bilgi görüntüden anlaşılamıyorsa bunu açıkça söyle. Yanıtı kısa, başlıklı maddeler hâlinde ver.This prompt asks for a description limited to observable elements. Narrow down the headings in the list to suit the content of the project image, as needed; do not ask the model to fill in information that is not shown.
2. Focus on a single design criterion. Visual question answering can be more targeted than asking for a general description. Limit the question to the design decision you want evaluated, and do not ask the model to infer information that is not in the image.
Bu iç mekân görselinde gün ışığının hangi bölgelere ulaştığını tarif et. Sadece görüntüde görülebilen ışık ve gölge dağılımına dayan. Pencere yönü veya günün saati görüntüden kesin olarak anlaşılamıyorsa tahmin yürütme; belirsizliği belirt.This example focuses on a single topic that can be examined in the image, such as the distribution of light and shadow. It also explicitly tells the model to leave out the direction or time of day if the image does not establish them.
3. Classify the surfaces visible in the image. The Gemini documentation covers classification tasks. Define the categories and output structure in the prompt. The aim of this example is to group surfaces observed in the image, not to verify technical standards.
Görselde seçilebilen yüzeyleri şu gruplara ayır: duvar, zemin, tavan ve sabit mobilya. Her grupta görülen yüzeyin rengini ve dokusunu yalnızca görseldeki izlenime göre tarif et. Malzeme türünü kesin olarak belirleyemiyorsan olası malzeme adı verme; “görselden belirlenemiyor” yaz.Predefined categories can make the results easier to review. Even so, it may not always be possible to identify the exact material from an image; stating this uncertainty in the prompt makes the limits of the response clear.
4. Describe spatial relationships in the image. Comparing multiple images in a single request is not shown in the source. So the example below focuses on spatial relationships visible in a single image.
Bu iç mekân görselinde mobilyaların ve dolaşım alanlarının kadraj içinde nasıl konumlandığını tarif et. Yalnızca görüntüde seçilebilen ilişkilere dayan; ölçü, geçiş genişliği veya planda görünmeyen başka bilgiler üretme. Emin olmadığın noktaları ayrıca belirt.This prompt does not ask the model to derive a plan or measurements; it asks only for relationships visible in the frame to be put into words. Treat such a description as an observation based on the image, not as a design decision.
5. Generate a short image description. If you need a short text for a presentation board, image archive, or internal team sharing, specify a word limit and content format instead of requesting a lengthy analysis. Check the response against the image before using it.
Bu mimari render için Türkçe, en fazla iki cümlelik tarafsız bir görsel açıklaması yaz. Mekânın kullanımını, kadrajda öne çıkan tasarım öğelerini ve görünen ışık koşullarını aktar. Görselde doğrulanamayan proje bilgisi, konum veya tasarımcı adı ekleme.This prompt defines a short, objective description format. Asking the model to leave out project information that cannot be verified in the image can also help prevent the text from being mistakenly used as project credits.
Common mistakes and checks
An open-ended instruction such as “Analyze this image” leaves it up to the model to decide what to examine and how to structure the response. Specify the task, observation criteria, and output format. Asking for a short list or specific headings can make the response easier to review.
Asking for information that is not visible in an image as if it were certain can also lead to misleading results. A render does not always show construction details or product names for materials. Add a constraint such as “state if this cannot be determined from the image” to the prompt, and compare important observations against the original image. Do not use the model response in place of project drawings, technical specifications, or an on-site inspection.
Choose the transfer method based on file size and whether you need to reuse the image. Inline image data has a 20 MB limit for the total request. Consider the Files API for large files or when you need to reuse an image. Check that the prompt and image are sent together and that the MIME type matches the image; the source examples also specify the image content type in the request.
Rather than treating the first response as final, refine the prompt. You can narrow the scope, clarify the output format, or specify how uncertainty should be expressed. Prompt design is an iterative process; it is recommended that you review the results of different attempts for your intended use.
How it fits into architecture and visualization workflows
This approach may be useful for teams that need to describe renders, classify elements in a single image, or organize observations ahead of a design meeting. Students can also generate draft descriptions for presentation images and check whether they are limited to what is actually visible.
However, the source documentation does not provide information about direct integration with architectural software, local hardware requirements, API pricing, or how project data is stored. Before adopting the workflow, separately review the cost and terms of the API access you use, office policy, and rules for sharing images. Check internal approvals before submitting images that are sensitive to a client or project.
Next steps
Start with a single non-critical image. Try description, classification, and visual question-answering prompts separately on the same image, then evaluate which response format meets your team's needs. For repeated analyses, consider using the Files API; for one-off, small requests, consider the appropriate transfer method.
If responses are inconsistent, make the prompt more specific: state what to examine, the limits of observation, and the desired output format. Compare the results against the image itself; do not treat information the model cannot see or the image cannot verify as project data.
Sources and license
This tutorial is adapted into Turkish from Google's Gemini API image understanding and prompt design documentation. Both sources are licensed under CC BY 4.0. The prompt examples for architectural images were created specifically for this guide.
Sources
2 sourcesSource texts are not republished; short quotes are marked, everything else is our own summary and commentary.
For architecture firms and visualization teams, the most practical use may be describing individual images and bringing observations into a consistent format for design reviews. This approach can support presentation preparation, but a model response is not a substitute for project validation or technical decisions.
For offices in Turkey, it is important to check API access costs, image-sharing policies, and terms regarding the use of client data before trying this approach. The sources do not explain pricing, hardware requirements, or data storage details. It would therefore be prudent to avoid starting with sensitive images and to compare outputs against the original image.
Frequently asked questions
How do you send an architectural image to the Gemini API?
You can send an image via an accessible URL, inline encoded data, or the Files API. The Files API is recommended for large files or images that will be reused.
What is the size limit for inline image data in the Gemini API?
The total request size, including text prompts and other content, must not exceed 20 MB. The Files API can be used for larger requests.
What image tasks can the Gemini API perform on architectural renders?
The Gemini documentation covers image description, classification, and visual question answering. Responses should be checked against what is actually visible in the image.



