מה העובד שלך יודע לפני שהוא בונה משהו

כל אחד יכול לכוון מודל לאתר אינטרנט. אלה החוקים שהוא חל על כל דף, למה כל אחד מהם קיים, ומי מהם לעצור פרסום במקום להזהיר אותו.

כל כלל להלן שולם על ידי תקרית באתר שאנחנו מנהלים □ צוק דירוג, תרגום שבור, דף שתרגם קוד מקור משלו. אנחנו לא מנחשים בתרגול הטוב ביותר; אנחנו כותבים מה שכבר השתבש באתר שלך.

חוקים שחוסמים פרסום באופן ישיר: 8 מ 12. דף חסום מועבר בחזרה לעובד עם הסיבה, והוא מנסה שוב. הוא לא מפרסם ומזהיר אותך לאחר מכן.

1. One page, one subject

בלוקים מפרסם

Exactly one <h1>. A descriptive <title> of 15-60 characters.

למה?: Title is the strongest per-page signal Google has. Google rewrites titles for the SERP most of the time, but the signal it forms from yours still affects ranking, and a generic or duplicated title reads as a quality problem.

2. A description written for a human

אזהרה

A meta description of 140-158 characters, unique to the page.

למה?: Description does not rank the page, it decides whether anyone clicks it. Too short wastes the snippet; too long truncates mid-sentence.

3. No shared body text, ever

בלוקים מפרסם

Every page gets body copy and FAQ answers written for that page. Never a shared template with a noun swapped in.

למה?: This is the one that costs months. A site we run shipped ~1700 pages on one generic FAQ body and a site-level quality classifier suppressed the whole domain — pages still indexed, ranking gone, no manual action to appeal. Recovery ran on core-update timescale.

4. Absolute canonical, paired with hreflang

בלוקים מפרסם

Every page self-canonicalises with an absolute URL. Any hreflang cluster ships alongside that canonical, also absolute.

למה?: Relative hrefs are silently dropped. hreflang without a self-canonical fills the index report with "Duplicate without user-selected canonical" — 56,801 pages on one site we run.

5. Markup may only claim what the page shows

בלוקים מפרסם

Structured data mirrors the visible text exactly. Schema goes on the page types that earn a rich result, not on everything.

למה?: Schema is not a general ranking input and not required for AI search -- Google says so directly. Markup that claims a price or a question the page never shows is a violation that can earn a manual action.

6. Django comments are {% comment %}

בלוקים מפרסם

Never {# #}. Not even on one line.

למה?: Django's template lexer regex has no DOTALL flag, so {# #} only closes on its own line. A multi-line one renders as visible page text, and if it contains a {% %} substring the lexer treats it as a real tag and every render crashes. Both have happened in production.

7. Positional {0} slots in translatable strings

בלוקים מפרסם

Any string that will be machine-translated uses {0}, {1} — never {name}.

למה?: MT engines translate or transliterate the word inside the braces, or drop a brace entirely. A numeric slot has no word to translate and survives intact. This silently broke about 10% of non-English rows on one site before anyone noticed.

8. Nothing is indexable until it is finished

בלוקים מפרסם

Pages ship noindex. Indexing is a deliberate opt-in by the owner, per site.

למה?: These sites share a wildcard domain. One spam tenant indexed on that wildcard damages every other tenant on it, and a half-finished page indexed on day one takes months to undo.

9. Build the thing, do not describe the thing

אזהרה

A page either does something useful or says something specific. A page that only describes a tool is not a page.

למה?: Thin wrappers around someone else's API are the clearest signal a quality classifier has. Real utility is what survives a core update.

11. Alt text and intrinsic dimensions on every image

אזהרה

Every <img> has alt text and explicit width/height.

למה?: Alt text is an accessibility requirement first and an indexing signal second. Missing dimensions cause layout shift, which is a Core Web Vitals failure on mobile — and mobile is what gets indexed.

12. Retired URLs return 410, verified

בלוקים מפרסם

A removed page returns 410 Gone, confirmed by an actual request after deploy.

למה?: A 404 lingers as a soft-404 for months; a 410 deindexes cleanly. And per-view dispatch often intercepts the request before the 410-returning view is ever reached, so the check has to be a real HTTP request, not a code read.

למה זה מוך ולא מדריך סגנון

כלל שחי רק בהוראה הוא הצעה ▪ המודל עוקב אחריו רוב הזמן ובשקט לא השאר, ואתה מגלה שלושה חודשים מאוחר יותר בדו"ח תנועה. התוצאה היא על הקבלה בכל מקרה, כך שאתה יכול לראות אילו בדיקות רצות. עמוד שלא מצליח כלל חסימה מוחזר עם הסיבה ונכתב. התוצאה היא על הקבלה בכל מקרה, כך שאתה יכול לראות שבדיקות רצות.

ושום דבר לא נמצא באינדקס עד שתגיד.

כל אתר מתחיל noindex. לא כברירת מחדל אתה צריך לגלות □ ככלל מאוכפים מוך, כי אתרים אלה חולקים תחום ושכן רע אחד פוגע בכולם על זה. כאשר האתר שלך שווה למצוא, אתה הופך אינדקס על בהגדרות וזה מתחיל להיות זחל.

הפעל המשמרת הראשונה שלך בחינם