# legalrag.lt — public case-law pages are indexable; the app itself is not. # Terms of use: https://app.legalrag.lt/salygos # Bulk/automated extraction of a substantial part of this database is prohibited # (EU Database Directive 96/9/EC sui generis right). # # POLICY: SEARCH + LIVE GROUNDING welcome, TRAINING not. # If our summaries are absorbed into model weights, answers quote us without # citing or linking. Crawlers that fetch us live and attribute get everything # public; crawlers that harvest for training get nothing. # # 2026-08-03: Cloudflare "Managed robots.txt" turned OFF and its block folded in # here by hand. Reason: it was all-or-nothing and disallowed Google-Extended. # Per Google's Vertex AI / Gemini docs, "Grounding with Google Search does not # use web pages for grounding that have disallowed Google-Extended" — i.e. the # managed block locked us out of grounding in Gemini apps and the Grounding # with Google Search product, an answer-engine channel we want. (AI Overviews # in Search itself is governed by ordinary Googlebot, not Google-Extended.) # Owning the file lets us keep the training reservation AND allow that grounding. # TRADE-OFF: new training crawlers are no longer added automatically — review # https://developers.google.com/crawling/docs/crawlers-fetchers and Cloudflare's # AI Crawl Control crawler list periodically. # # ⚠ robots.txt has NO INHERITANCE: a crawler obeys ONLY the most specific group # matching its name and ignores `User-agent: *` entirely. So every named group # below repeats the full rule set — a short block would leave that agent free to # crawl /api/, /login, /liteko and the saltiniai pages. # ── Content signals (contentsignals.org) ───────────────────────────────────── # As a condition of accessing this website, you agree to abide by the following # content signals: # (a) Content-Signal = yes → you may collect content for that use. # (b) Content-Signal = no → you may not collect content for that use. # (c) absent signal → neither granted nor restricted via Content-Signal. # Meanings: # search: building a search index and returning hyperlinks/short excerpts. # ai-input: inputting content into AI models (RAG, grounding, real-time taking # of content for generative AI answers). # ai-train: training or fine-tuning AI models. # use: how AI systems may consume the content (immediate/reference/full). # # ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF # RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT # AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET. User-agent: * Content-Signal: search=yes,ai-input=yes,ai-train=no,use=reference Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai # ── Training crawlers: no ──────────────────────────────────────────────────── # These exist to build training corpora, not to send traffic or cite sources. # NB the search/answer sibling of each is allowed further below: ClaudeBot is # blocked but Claude-SearchBot/Claude-User are not; Applebot-Extended is blocked # but Applebot is not; Google-Extended is ALLOWED (it gates grounding, and Google # Search grounding honours it — see the note at the top). User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: meta-externalagent Disallow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: CloudflareBrowserRenderingCrawler Disallow: / # ── Search + answer engines: yes (fetch live, attribute) ───────────────────── User-agent: Googlebot Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai # Gates Gemini apps AND Grounding with Google Search. Allowed deliberately — # same treatment as every other live-grounding answer engine below. User-agent: Google-Extended Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai User-agent: bingbot Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai User-agent: OAI-SearchBot Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai User-agent: ChatGPT-User Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai User-agent: Claude-SearchBot Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai User-agent: Claude-User Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai User-agent: PerplexityBot Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai User-agent: DuckAssistBot Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai User-agent: Applebot Allow: /byla/ Allow: /teisejai Allow: /salygos Allow: /$ Disallow: /app/ Disallow: /api/ Allow: /api/sitemaps/ Disallow: /login Disallow: /auth/ Disallow: /liteko Disallow: /teisejai/*/saltiniai Sitemap: https://app.legalrag.lt/sitemap.xml