# ============================================================================ # Content Signals Policy — https://contentsignals.org # ============================================================================ # # NOTICE: any use of this site's content is subject to the Content-Signal # directives declared in the groups below. As used here: # # search means building a search index and providing search results # (e.g. returning hyperlinks and short excerpts). It does NOT # include providing AI-generated search summaries. # ai-input means inputting content into one or more AI models (e.g. # retrieval augmented generation, grounding, or other real-time # use of content to produce generative AI answers). # ai-train means training or fine-tuning AI models. # # "yes" is express permission for that use. "no" is the absence of express # permission. An omitted signal expresses no preference. # # Signals are scoped per user-agent group (RFC 9309: a crawler obeys only the # most specific group that matches it), so the two groups in this file say # deliberately different things: # # User-agent: * → search + AI answers welcome, training not permitted. # Reaches Googlebot, Bingbot, OAI-SearchBot, # ChatGPT-User, Claude-User, PerplexityBot, CCBot, # Bytespider, Meta-ExternalAgent, et al. # AI training funnel → training expressly permitted, because that group's # crawlable surface is a curated set of public # marketing/blog/job-board pages (Disallow: / # fallback). See its own header below. # # These signals are advisory. Access control is enforced by the Allow/Disallow # rules here plus Cloudflare bot rules — not by this policy. User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: /candidates/signup Allow: /pt/candidates/signup Allow: /en/candidates/signup Allow: /es/candidates/signup Allow: /candidates/faq Allow: /pt/candidates/faq Allow: /en/candidates/faq Allow: /es/candidates/faq Disallow: /candidates Disallow: /pt/candidates Disallow: /en/candidates Disallow: /es/candidates Disallow: /company_users Disallow: /pt/company_users Disallow: /en/company_users Disallow: /es/company_users Disallow: /companies Disallow: /pt/companies Disallow: /en/companies Disallow: /es/companies Disallow: /vagas- Disallow: /pt/vagas- Disallow: /en/vagas- Disallow: /es/vagas- # Non-canonical vagas/jobs page URLs Disallow: /jobs Disallow: /pt/jobs Disallow: /en/vagas Disallow: /es/vagas # Login/Sign in pages Disallow: /entrar Disallow: /pt/entrar Disallow: /en/entrar Disallow: /es/entrar # Legacy Data privacy policy Disallow: /data-privacy-policy Disallow: /pt/data-privacy-policy Disallow: /en/data-privacy-policy Disallow: /es/data-privacy-policy # Non-canonical privacy policy URLs (canonical: /politica-de-privacidade, /en/privacy-policy, /es/politica-de-privacidad) Disallow: /privacy-policy Disallow: /politica-de-privacidad Disallow: /en/politica-de-privacidade Disallow: /en/politica-de-privacidad Disallow: /es/politica-de-privacidade Disallow: /es/privacy-policy # Non-canonical terms of use URLs (canonical: /termos-de-uso, /en/terms-of-use, /es/condiciones-de-uso) Disallow: /terms-of-use Disallow: /condiciones-de-uso Disallow: /en/termos-de-uso Disallow: /en/condiciones-de-uso Disallow: /es/termos-de-uso Disallow: /es/terms-of-use # Marketing pages without locale prefix. # # The following paths 301-redirect to their /en/ canonical at Cloudflare (verified # 2026-07-19). They are intentionally left crawlable: a bot must be allowed to fetch # the bare path to observe the 301 and consolidate onto /en/. Blocking them here would # hide the redirect, so Wix's hreflang references would keep the unprefixed URLs stuck # as "known but blocked" (URL-only indexing) — the problem this section used to cause: # /ats /tech-recruiter /talent-pool /pricing /case-of-success /cases-of-success # /faq /the-geek-advantage /build-my-own-team /job-board /autopilot /complete-solution # /glossary /glossary-index /post /blog /privacy-portal /privacy-policies # # The paths below still serve 200 with NO redirect, so they stay blocked to avoid # duplicate content with the /en/ canonical. Disallow: /login Disallow: /thank-you-page-english-form Disallow: /thank-you-page-spanish-form # Machine-readable partner ingestion endpoints — not pages, and not a surface # for search indexing or AI training. The same jobs are crawlable as HTML at # /:locale/:company/jobs/:slug (JobPosting structured data), which is the # surface meant for discovery. # # Not restated in the AI-training group below: its `Disallow: /` fallback # already blocks both, and neither path appears in that group's Allow list. # Rationale + access-control details are internal (see the WAF feeds runbook). Disallow: /api/ Disallow: /feeds/ # Cloudflare internal URLs — not real pages # Allow image-resizing endpoint so crawlers can fetch optimised images Allow: /cdn-cgi/image/ Disallow: /cdn-cgi/ # ============================================================================ # Custom AI Training Funnel (Allows specific public assets, locks out the rest) # ============================================================================ # # ai-train=yes is intentional here. The crawlable surface of this group is only # the public marketing collateral allowlisted below (Disallow: / fallback at the # end) — brand pages, blog, glossary, cases, public job boards. Training on that # corpus is free brand distribution, not a licensable asset. Everything with # licensing value (candidate profiles, company data, internal apps) is blocked # for every user-agent, including this group's Disallow: / fallback. # # search/ai-input are declared for completeness; the four user-agents below are # training-only crawlers (their search and live-fetch counterparts — OAI-SearchBot, # ChatGPT-User, Claude-User — match User-agent: * instead and read its signals). # # Reversibility note: flipping ai-train to "no" later governs future crawls only; # it does not untrain existing models, and content already in Common Crawl keeps # propagating. Actual enforcement of a future opt-out means removing these Allow # rules or enabling Cloudflare pay-per-crawl, not editing this value. User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended Content-Signal: search=yes, ai-input=yes, ai-train=yes # The curated content map itself. Without this, the Disallow: / fallback below hides # /llms.txt from exactly the crawlers it is meant to guide. Allow: /llms.txt # Allow exactly the root homepages per localized subfolder Allow: /$ Allow: /pt/$ Allow: /en/$ Allow: /es/$ Allow: /pt$ Allow: /en$ Allow: /es$ # PT Marketing Pages Allow: /pt/empresas/* Allow: /pt/tech-recruiter-as-a-service Allow: /pt/contratar-programador Allow: /pt/report-salarial Allow: /pt/tour-interativo Allow: /pt/tour-virtual-plano-lite Allow: /pt/talent-recruiter-as-a-service Allow: /pt/contratar-talentos # PT case studies live under /pt/empresas/casos-de-sucesso*, already covered by the # /pt/empresas/* wildcard above. There is no top-level /pt/casos-de-sucesso — it 404s # on both hosts (checked 2026-08-07), so do not re-add an Allow for it. Allow: /pt/blog Allow: /pt/blog/* # English (EN) Marketing Pages Allow: /en/ats Allow: /en/ai-recruiting Allow: /en/tech-recruiter Allow: /en/talent-pool Allow: /en/pricing Allow: /en/case-of-success Allow: /en/cases-of-success/* Allow: /en/faq Allow: /en/the-geek-advantage Allow: /en/build-my-own-team Allow: /en/job-board Allow: /en/autopilot Allow: /en/complete-solution Allow: /en/glossary-index Allow: /en/glossary/* Allow: /en/blog Allow: /en/post/* # Spanish (ES) Marketing Pages Allow: /es/ats Allow: /es/ai-recruiting Allow: /es/tech-recruiter Allow: /es/talent-pool Allow: /es/pricing Allow: /es/case-of-success Allow: /es/cases-of-success/* Allow: /es/faq Allow: /es/the-geek-advantage Allow: /es/build-my-own-team Allow: /es/job-board Allow: /es/autopilot Allow: /es/complete-solution Allow: /es/glossary-index Allow: /es/glossary/* Allow: /es/blog Allow: /es/post/* # Allow Multi-lingual Company Job Boards & Individual Postings Allow: /pt/*/jobs Allow: /pt/*/jobs/* Allow: /en/*/jobs Allow: /en/*/jobs/* Allow: /es/*/jobs Allow: /es/*/jobs/* Allow: /pt/vagas Allow: /en/jobs Allow: /es/jobs Allow: /candidates/faq Allow: /pt/candidates/faq Allow: /en/candidates/faq Allow: /es/candidates/faq # Allow the static assets behind the WPEngine-backed /pt marketing pages. # These are served from root /wp-content on this host, not from under the page paths # above, so the Allow rules for /pt/empresas/* etc. do not reach them and the # Disallow: / fallback below would block every image on those pages. Scoped to # /uploads because that is where the images are; plugin and theme code is of no use to # a crawler that does not render. No matching rule is needed in the User-agent: * # group — it has no blanket Disallow, so these paths are already crawlable there. Allow: /wp-content/uploads/ # Enforce a strict fallback block to protect all internal routes, candidate profiles, and apps Disallow: / Sitemap: https://www.geekhunter.com/sitemap.xml