Latest news
Distribute Architectural Elements with Blender Geometry NodesSampling Along Curves with Blender Geometry Nodes: A Sample Curve GuideUnreal Engine Movie Render Queue Guide: Architectural Rendering SettingsBlender Asset Browser Guide: Build an Asset Library for ArchitectureAnalyze Architectural Videos with the Gemini API: A Step-by-Step GuideGetting ready for the Archicad–Autodesk Forma connection: a guide to seamless BIM data transfer4 settings that fix wall junction errors in Archicad: composite wall guidePhasing Management in Revit: A Step-by-Step Guide for Renovation ProjectsHow to create Dynamic Components in SketchUp: smart components with formulas3ds Max'ta Shell, Kamera ve MassFX Rigid Body ile Sunum Sahnesi Hazırlama Rehberi
TutorialsGoogleGemini Omni FlashVeo 3.1

Analyze Architectural Videos with the Gemini API: A Step-by-Step Guide

The Gemini API can summarize videos, answer questions about their content, and provide information tied to timestamps. This guide explains how to process architectural and interior videos according to file size and verify the results.

A contemporary interior tour in natural daylight evokes the theme of video analysisAI image
Representative image, generated with AI.Image: 3dsınıfı / FCA AI

In brief

  1. Gemini models can describe videos, answer questions about their content, and refer to specific timestamps.
  2. Google recommends the Files API for files of 100 MB or more and for videos you plan to reuse.
  3. The documentation gives both an under-100-MB limit for inline video and an under-20-MB limit for the total request; following the lower limit in the implementation guidance is the cautious choice.
  4. Gemini Omni Flash and Veo 3.1 are options for video generation; the example in this guide to analyzing existing videos uses the gemini-3.8-flash model.

What will you learn in this guide?

This tutorial walks you through a basic workflow for analyzing an architectural presentation, construction-site recording, interior tour, or design-process video with the Gemini API. The goal is not to generate a new video, but to submit an existing one to the model and ask for a summary, answers about specific scenes, or a content breakdown. Google’s video understanding documentation lists the models’ ability to describe and segment videos, extract information, answer questions about video content, and refer to timestamps.

The examples include text prompts to send to the API. They are not commands that guarantee a particular result; they help you specify what kind of information you want from the video. Before using the output for a project decision, measurement, or technical verification, compare it with the source footage.

Requirements and choosing a file

The workflow requires access to the Gemini API and an integration that can submit the video. The official examples use Python, JavaScript, Java, Go, and REST; the steps below are not tied to a particular programming language. The REST example uses a key called GEMINI_API_KEY. The video documentation used here does not explain API access or account requirements in detail.

The example API request specifies the model name gemini-3.8-flash. This is the model named in the video understanding example. There is a separate guide for video generation: Google mentions Gemini Omni Flash for general video generation and dialogue-based editing, and Veo 3.1 for needs such as extending a scene or controlling the final frame. Don’t confuse these two use cases.

Before uploading, check the file size and whether you’ll reuse it:

  • Google recommends the Files API for videos over 100 MB, long videos, or files you’ll reuse. The table in the documentation lists Files API limits of 2 GB for free usage and 20 GB for paid usage.

  • The Cloud Storage Registration option specifies a 2 GB limit per file and says there is no additional storage limit.

  • The YouTube URL method is listed for public YouTube videos.

  • Inline data is suitable for small, one-off files. The documentation table associates this method with videos under 100 MB and short clips; the implementation section also says the total request must be under 20 MB. Since these are different thresholds, following the lower 20 MB total-request limit is the cautious approach for inline uploads.

Step-by-step video analysis workflow

  1. Decide what you want to find out from the video. Instead of a general request to “analyze the video,” describe the information you’re looking for. For example, an interior tour could be broken down into room layout, visible materials, and the areas the camera passes through. Don’t assume the model can identify measurements or concealed structural layers that aren’t visible in the footage.

  2. Upload the file using the appropriate method. For repeated use or large files, choose the Files API. In the official example workflow, the video is uploaded, processing is allowed to finish, and the file reference is sent with the text request once its status is ACTIVE. If processing fails, the documentation example treats the FAILED status as an error. Inline data can be used for smaller videos; check the file size against the total-request limit.

  3. Send the request with the video content. The request includes the video input and the analysis prompt as text. In the Python example, the file is uploaded through the Files API; the example for smaller content sends the video bytes inline. Whichever integration you use, make sure the video MIME type and file reference are passed correctly.

  4. Ask for a limited output on the first try. You can add the following examples separately to the text portion of your request. Each is geared toward a different architectural review task.

Prompt
Bu mimari iç mekân videosunu sahne sırasına göre özetle. Görüntüde açıkça görülen mekânları, malzeme izlenimlerini ve kamera hareketlerini ayrı maddeler hâlinde belirt. Görüntüde olmayan bilgileri tahmin etme; belirsiz noktaları açıkça işaretle.

This prompt separates the general description into spaces, materials, and camera movement. Treat descriptions of materials as visual impressions from the video, not as definitive technical specifications.

Prompt
Videoda girişten ana yaşam alanına geçişin görüldüğü bölümü bul. İlgili zaman damgasını belirt ve o anda kadrajda görünen mekânsal ilişkileri kısaca açıkla. Bölüm net değilse bunu söyle; zaman veya plan bilgisi uydurma.

This prompt asks about a scene by timestamp. Open the video and check the time given in the model’s response. Although the documentation says the model can refer to timestamps, it does not guarantee a perfect match in every request.

Prompt
Bu şantiye videosunda tamamlanmış görünen imalatları ve hâlen devam ettiği açıkça anlaşılan işleri ayır. Her bulguyu videoda görülen sahneyle ilişkilendir; yapısal güvenlik, ölçü veya mevzuat uygunluğu hakkında çıkarım yapma.

This example asks for a classification that can help organize visual records; it does not produce a safety or compliance assessment from the video. Compare the response with the site team’s records.

  1. Review the response and narrow the prompt. If the summary is too general, ask a new question describing a specific space, scene, or time range shown in the video. Gemini’s video understanding capabilities include asking questions about content and referring to timestamps. Verify important findings in the video before using the result.

Common mistakes

Not checking the file-size requirements: The Files API is recommended for large files or files you’ll reuse. Even if you’re considering the inline method for a small file, the documentation also specifies a 20 MB limit for the total request; checking only the video file’s size may not be enough.

Requesting analysis before processing is complete: The Files API example waits after upload until the file is ACTIVE. Using the file reference before checking its processing status does not follow the example workflow.

Confusing video analysis with video generation: Use the video understanding guide to review an existing recording. The documentation describes Gemini Omni Flash and Veo 3.1 as separate workflows for video generation and editing.

Drawing technical conclusions from visuals: Ask the model to describe visible elements; don’t present measurements, material properties, or construction compliance that aren’t shown in the footage as verified facts. Check timestamp and scene matches against the source video, too.

Implications for architecture and interior design workflows

This method can be a starting point for architecture firms, interior designers, and students who want to find scenes in long presentation videos, organize an interior tour by topic, or identify visually observable work in a site recording. The Files API approach is practical, particularly for large files you’ll reuse; inline uploads are an alternative for small, one-off tests. However, the source does not explain usage costs or specific hardware requirements. Costs and compatibility should therefore be checked separately before implementation, and model responses should not be treated as technical approval.

Next steps

Start with a short, representative recording and request a general summary; then narrow the output with questions about scenes or timestamps. If you’ll submit the same video more than once, consider Google’s recommendation to use the Files API. If you also need video generation in the API, review the Gemini Omni Flash and Veo 3.1 documentation separately from the video understanding workflow.

Sources and license

This guide has been adapted into Turkish based on information in Google’s Gemini API video understanding and video generation documentation. Both sources are licensed under CC BY 4.0.

Sources

2 sources
A(
ai.google.dev (CC BY 4.0)ai.google.dev/gemini-api/docs/video-understanding
Summary
A(
ai.google.dev (CC BY 4.0)ai.google.dev/gemini-api/docs/video
Summary

Source texts are not republished; short quotes are marked, everything else is our own summary and commentary.

3dsınıfı’s take
3dEditor’s assessment

For architecture firms and students, the most useful application is speeding up the process of finding scenes in long videos and producing an initial summary. However, a model’s output is an interpretation of the footage; it cannot replace verification of measurements, material specifications, or construction-site compliance.

Teams in Turkey should start with a short sample video and narrowly scoped questions. Since the documentation does not specify pricing or hardware requirements, costs and API access conditions should be checked separately. File size and frequency of reuse will also determine whether to choose the Files API or inline uploads.

Frequently asked questions

Can the Gemini API analyze architectural videos?

Gemini models can describe videos, answer questions about their content, and refer to specific timestamps. Technical conclusions should be verified against the source footage.

How do you upload large videos to the Gemini API?

Google recommends the Files API for files of 100 MB or more and for files you plan to reuse. The documentation lists Files API limits of 2 GB for free usage and 20 GB for paid usage.

Which model is used in the video understanding example?

The example in Google’s video understanding documentation uses the `gemini-3.8-flash` model. The documentation describes Gemini Omni Flash and Veo 3.1 as separate options for video generation.

Comments and the forum are in Turkish.Join the discussion
+

Related news