Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 15 additions & 5 deletions docs/news-operations.md
Original file line number Diff line number Diff line change
Expand Up @@ -149,17 +149,19 @@ read-only research process. It consumes subscription usage. Paid API variables
are removed and there is no paid fallback. Publication occurs only in the local
validated runner; the research process has no publication credentials.

The publisher explicitly selects `gpt-5.6-sol` with medium reasoning for research
and Russian editorial translation, and low reasoning for social market matching.
The publisher explicitly selects the lower-cost `gpt-5.6-luna`. Research keeps
medium reasoning behind the unchanged source, evidence, novelty and cover gates;
Russian editorial translation and social market matching use low reasoning.
Only research enables live web search. Host skill discovery, plugins, account
connectors, shell access and agent delegation are disabled for these invocations;
shared Codex configuration is not modified. Successful invocations append token
counts, cached input, purpose and duration timestamps to private
`writer-usage.jsonl` files without storing article text or credentials there.
A provider usage-limit refusal stops the research round immediately and defers
the next attempt for thirty minutes while retaining the staged edition. Timer
checks during that pause make no model calls; one refusal ends the next attempt
immediately. There is no paid fallback or automatic maintenance-agent wakeup.
the next attempt for 6 hours, then 12 hours and at most 24 hours while retaining
the staged edition. Timer checks during that pause make no model calls; one
refusal ends the next attempt immediately. There is no paid fallback or
automatic maintenance-agent wakeup.
The five-hour selector makes no model call when no market candidates exist; it uses
the existing verified-article country/topic ordering and publishes news only.
When matching is needed, article context is supplied once per article and market
Expand All @@ -178,6 +180,14 @@ from this host; it is not bypassed. Model and private caches are excluded from G
The separate Python environment is `/root/OddsFront/.local/translation-venv`;
the model is `/root/OddsFront/.local/translation-model`.

English publication sends IndexNow and WebSub immediately, then queues
`oddsfront-news-translation.service` without waiting for it. The isolated local
worker may use two CPU cores, keeps the same private cache, and translates all
text needed to finish the newest articles before using a later pass for archive
or map-dictionary backlog. It exports only complete per-language articles, then
sends another IndexNow delta and WebSub notification for the newly available
localized URLs. Translation never delays the next English sitemap update.

Pinned packages: ctranslate2 4.8.2, transformers 4.57.6, sentencepiece 0.2.1,
PyTorch 2.7.1+cpu. A model cache hit does not repeat inference. Translations are
marked as machine translations and link to English. Empty or over-budget text
Expand Down
5 changes: 5 additions & 0 deletions docs/search-indexing.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,11 @@ URL, reciprocal `hreflang` entries, `x-default`, localized social metadata, and
- `/sitemap.xml` links to `/sitemaps/core.xml` and bounded article sitemap
parts. Withdrawn articles are excluded. The core sitemap includes the
crawlable news archive, archive pagination, and the newsdesk page.
- General sitemaps list one high-value English `<loc>` per article, listing,
topic and country page, with every complete localized version attached as a
reciprocal `hreflang` alternate. HTML pages retain the same reciprocal links.
This keeps the submitted crawl inventory focused instead of multiplying it by
every language, while search engines can still discover and serve each locale.
- `/news-sitemap.xml` is the Google News sitemap index. It splits the last 48
hours by language and then into parts of no more than 1,000 news URLs under
`/news-sitemaps/{locale}-{part}.xml`. Simplified Chinese uses Google News's
Expand Down
18 changes: 8 additions & 10 deletions lib/news/sitemap.ts
Original file line number Diff line number Diff line change
Expand Up @@ -60,11 +60,9 @@ export function coreSitemap(catalog: NewsCatalog): string {
updatedAt: catalog.updatedAt,
priority: "0.6",
})),
...LOCALES.flatMap((locale) => [
urlEntry({ path: newsPath(locale), updatedAt: catalog.updatedAt, priority: "0.9", alternateLinks: alternates("/news") }),
...categories.map(topic=>urlEntry({ path:newsTopicPath(topic,locale), updatedAt:catalog.updatedAt, priority:"0.8", alternateLinks:alternates(newsTopicPath(topic,"en")) })),
...countries.map((country) => urlEntry({ path: newsCountryPath(country, locale), updatedAt: catalog.updatedAt, priority: "0.8", alternateLinks: alternates(`/news/${country}`) })),
]),
urlEntry({ path: newsPath("en"), updatedAt: catalog.updatedAt, priority: "0.9", alternateLinks: alternates("/news") }),
...categories.map(topic=>urlEntry({ path:newsTopicPath(topic,"en"), updatedAt:catalog.updatedAt, priority:"0.8", alternateLinks:alternates(newsTopicPath(topic,"en")) })),
...countries.map((country) => urlEntry({ path: newsCountryPath(country, "en"), updatedAt: catalog.updatedAt, priority: "0.8", alternateLinks: alternates(`/news/${country}`) })),
];
return sitemapDocument(urls.join(""));
}
Expand All @@ -73,16 +71,16 @@ export function articleSitemap(catalog: NewsCatalog, page: number): string | nul
const activeArticles = catalog.articles.filter((article) => !article.withdrawal);
const offset = (page - 1) * SITEMAP_ARTICLE_BATCH_SIZE;
if (!Number.isInteger(page) || page < 1 || (offset >= activeArticles.length && page !== 1)) return null;
const urls = activeArticles.slice(offset, offset + SITEMAP_ARTICLE_BATCH_SIZE).flatMap((article) => {
const urls = activeArticles.slice(offset, offset + SITEMAP_ARTICLE_BATCH_SIZE).map((article) => {
const locales = articleLocales(article);
Comment on lines +74 to 75

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep every hreflang variant as a sitemap URL

For sitemap-based hreflang, each alternate URL must have its own <url><loc> entry and repeat the complete alternate set, including itself. Mapping every article to only the English URL here—and doing the same for listing, topic, and country pages above—leaves the localized href values without reciprocal sitemap entries, so crawlers consuming the sitemap annotations cannot validate those language clusters and the localized pages are no longer directly submitted in the general sitemap.

Useful? React with 👍 / 👎.

const alternateLinks = alternates(newsArticlePath(article, "en"), locales);
return locales.map((locale) => urlEntry({
path: newsArticlePath(article, locale),
return urlEntry({
path: newsArticlePath(article, "en"),
updatedAt: article.updatedAt,
priority: "0.8",
alternateLinks,
image: `${ORIGIN}/social/news/${locale.toLowerCase()}/${article.slug}?v=${encodeURIComponent(article.updatedAt)}`,
}));
image: `${ORIGIN}/social/news/en/${article.slug}?v=${encodeURIComponent(article.updatedAt)}`,
});
});
return sitemapDocument(urls.join(""));
}
14 changes: 10 additions & 4 deletions lib/news/writer.ts
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,11 @@ export type WriterUsage = { inputTokens: number; cachedInputTokens: number; outp
export type CodexProcessRunner = (input: { invocation: CodexWriterInvocation; prompt: string; cwd: string; timeoutMs: number }) => Promise<WriterUsage | void>;
export class CodexUsageLimitError extends Error {}
const PAID_API_ENV_KEYS = new Set(["OPENAI_API_KEY", "OPENAI_BASE_URL", "OPENAI_ORG_ID", "OPENAI_ORGANIZATION", "OPENAI_PROJECT_ID"]);
export const NEWS_WRITER_MODEL = "gpt-5.6-luna";

function reasoningEffort(purpose: WriterPurpose | undefined) {
return !purpose || purpose === "research" ? "medium" : "low";
}
export function createCodexWriterInvocation(input: {
schemaPath: string;
outputPath: string;
Expand All @@ -25,8 +30,8 @@ export function createCodexWriterInvocation(input: {
"exec",
"--ephemeral",
"--ignore-user-config",
"--model", input.model || "gpt-5.6-sol",
"-c", `model_reasoning_effort="${input.purpose === "selection" ? "low" : "medium"}"`,
"--model", input.model || NEWS_WRITER_MODEL,
"-c", `model_reasoning_effort="${reasoningEffort(input.purpose)}"`,
// A news editor needs public search, not hundreds of host skills, account
// connectors, shell tools or other agents. Do not change shared host config.
"--enable", "skip_host_skill_discovery",
Expand Down Expand Up @@ -111,12 +116,13 @@ export async function executeSubscriptionCodex(input: {
const schemaPath = path.join(workdir, "article.schema.json");
const outputPath = path.join(workdir, "article.json");
const startedAt = new Date().toISOString();
const model = input.model || NEWS_WRITER_MODEL;
try {
await writeFile(schemaPath, JSON.stringify(input.schema), "utf8");
const invocation = createCodexWriterInvocation({
schemaPath,
outputPath,
model: input.model,
model,
purpose: input.purpose,
env: input.env,
});
Expand All @@ -126,7 +132,7 @@ export async function executeSubscriptionCodex(input: {
cwd: workdir,
timeoutMs: input.timeoutMs || 240_000,
});
if (input.usageFile) await appendFile(input.usageFile, `${JSON.stringify({ startedAt, finishedAt: new Date().toISOString(), purpose: input.purpose || "research", model: input.model || "gpt-5.6-sol", promptBytes: Buffer.byteLength(input.prompt), ...usage })}\n`, { mode: 0o600 }).catch(() => {
if (input.usageFile) await appendFile(input.usageFile, `${JSON.stringify({ startedAt, finishedAt: new Date().toISOString(), purpose: input.purpose || "research", model, promptBytes: Buffer.byteLength(input.prompt), ...usage })}\n`, { mode: 0o600 }).catch(() => {
console.warn("Writer usage logging unavailable; retaining the completed editorial result.");
});
return await readFile(outputPath, "utf8");
Expand Down
26 changes: 26 additions & 0 deletions ops/systemd/oddsfront-news-translation.service
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
[Unit]
Description=OddsFront asynchronous local news translation worker
After=network-online.target oddsfront-news.service
Wants=network-online.target
ConditionPathExists=/root/OddsFront/scripts/news/translate.py

[Service]
Type=oneshot
WorkingDirectory=/root/OddsFront
ExecStart=/root/OddsFront/.local/translation-venv/bin/python scripts/news/translate.py
ExecStartPost=-/usr/local/bin/node scripts/news/indexnow.mjs
ExecStartPost=-/usr/local/bin/node scripts/news/websub.mjs
Comment on lines +11 to +12

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Notify search services after partial translation runs

If the bounded worker times out at 90 minutes or fails after exporting one or more languages, systemd does not execute ExecStartPost; the leading - only ignores failures of the post command itself. Consequently, successfully exported localized URLs from a partial run receive neither the promised IndexNow delta nor WebSub notification, and repeated late-language failures can prevent notifications indefinitely; place these notifications in an always-run cleanup path such as ExecStopPost or a wrapper.

Useful? React with 👍 / 👎.

Environment=NODE_ENV=production
Environment=PATH=/usr/local/bin:/usr/bin:/bin
Nice=10
CPUQuota=200%
MemoryMax=2G
TasksMax=128
PrivateTmp=true
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=read-only
ReadWritePaths=/root/OddsFront/.local/news /opt/oddsfront-market-feed/news /root/.codex
InaccessiblePaths=-/opt/arctrenches -/opt/arctrenches-legacy-6c13c07 -/root/DropsAnalytics -/root/DropsAnalytics-worktrees -/opt/sunder -/root/.ssh -/etc/oddsfront-market-feed.env
UMask=0077
TimeoutStartSec=90min
2 changes: 1 addition & 1 deletion scripts/news/run-edition.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -212,9 +212,9 @@ if (remaining.length) {
console.log(JSON.stringify(receipt));
if (process.argv.includes("--with-followups")) {
for (const [command, args, timeout] of [
["/root/OddsFront/.local/translation-venv/bin/python", ["scripts/news/translate.py"], 55 * 60_000],
[process.execPath, ["scripts/news/indexnow.mjs"], 60_000],
[process.execPath, ["scripts/news/websub.mjs"], 60_000],
["systemctl", ["start", "--no-block", "oddsfront-news-translation.service"], 30_000],
]) {
const result = spawnSync(command, args, { stdio: "inherit", env: process.env, timeout });
if (result.status !== 0) console.error(JSON.stringify({ status: "followup-failed", command: path.basename(command), exitCode: result.status }));
Expand Down
9 changes: 7 additions & 2 deletions scripts/news/translate.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,11 @@ def select_pending_texts(active_articles, dictionary_texts, language, cache, cac
article_pending.update(missing)
if len(article_pending) >= budget:
break
# Fresh, complete article translations are the search and readership
# priority. Do not delay their per-language export by filling the same
# inference pass with map labels or archive dictionary work.
if article_pending:
return sorted(article_pending)
Comment on lines +47 to +48

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Schedule the deferred dictionary translation pass

When regular editions continue publishing, every worker invocation begins with missing article text, so this early return discards the remaining text budget and dictionary_pending is never reached. Because the worker is only queued after a new edition, live market labels and explanatory-page strings can remain untranslated indefinitely even when the fresh articles consume far less than the 150-text budget; run a second pass after exporting the article translations or explicitly schedule the deferred backlog work.

Useful? React with 👍 / 👎.

remaining = max(0, budget - len(article_pending))
dictionary_pending = [
text for text in sorted(dictionary_texts - article_pending)
Expand Down Expand Up @@ -170,8 +175,8 @@ def backlog(language):
owners.append(text)
translated = {text: [] for text in pending}
expected_pieces = Counter(owners)
for start in range(0, len(pieces), 16):
batch = pieces[start:start+16]
for start in range(0, len(pieces), 32):
batch = pieces[start:start+32]
results = translator.translate_batch(batch, target_prefix=[[f"__{target}__"] for _ in batch], beam_size=2, max_batch_size=512, batch_type="tokens", max_decoding_length=320, max_input_length=512, no_repeat_ngram_size=4)
for index, result in enumerate(results):
tokens = [token for token in result.hypotheses[0] if not token.startswith("__") and token not in ["</s>", "<s>"]]
Expand Down
8 changes: 7 additions & 1 deletion tests/news-publication.spec.ts
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ import { NEWS_MAX_RESEARCH_ROUNDS, NEWS_RESEARCH_BATCH_SIZE, isPublishableEditio
import { buildIndexNowPlan, buildIndexNowSnapshot, indexNowBatches } from "../lib/news/indexnow";
import { NEWS_SITEMAP_URL_LIMIT, newsSitemap, newsSitemapIndex } from "../lib/news/news-sitemap";
import { newsRss } from "../lib/news/rss";
import { articleSitemap, coreSitemap } from "../lib/news/sitemap";
import { articleSitemap, coreSitemap, SITEMAP_ARTICLE_BATCH_SIZE } from "../lib/news/sitemap";

function draft():NewsDraft {
return {publishable:true,rejectionReason:"",alert:{eligible:false,kind:"none",actorCountries:[],targetCountries:[]},title:"Test fixture: regional diplomatic review",description:"Development-only publication validation fixture.",countries:["UA"],topics:["diplomacy"],
Expand Down Expand Up @@ -56,8 +56,14 @@ test("search discovery stays bounded, delta-based, and excludes withdrawn URLs",
const core = coreSitemap(catalog);
expect(core).toContain("https://oddsfront.com/news/about");
expect(core).toContain("https://oddsfront.com/news/archive?page=2");
expect(core).toContain('hreflang="ru" href="https://oddsfront.com/ru/news"');
expect(core).not.toContain("<loc>https://oddsfront.com/ru/news</loc>");
expect(core).not.toContain("https://oddsfront.com/global-conflict-map");
expect(core).not.toContain("https://oddsfront.com/news/world");
const articlePart = articleSitemap(catalog, 1)!;
expect(articlePart.match(/<url>/g)).toHaveLength(SITEMAP_ARTICLE_BATCH_SIZE);
expect(articlePart).toContain('hreflang="ru"');
expect(articlePart).not.toContain("<loc>https://oddsfront.com/ru/news/");
expect(newsRss(catalog, "en")).toContain('<atom:link href="https://pubsubhubbub.appspot.com/" rel="hub"/>');

const initialCatalog = { ...catalog, articles: articles.slice(0, 2) } as NewsCatalog;
Expand Down
8 changes: 6 additions & 2 deletions tests/news-writer.spec.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,20 +3,24 @@ import { mkdtemp, mkdir, readFile, rm, writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { spawnSync } from "node:child_process";
import { createCodexWriterInvocation, executeSubscriptionCodex } from "../lib/news/writer";
import { createCodexWriterInvocation, executeSubscriptionCodex, NEWS_WRITER_MODEL } from "../lib/news/writer";

test("subscription research keeps search but cannot load host accounts or paid API credentials", () => {
const invocation = createCodexWriterInvocation({ schemaPath: "schema.json", outputPath: "article.json", env: { PATH: "/usr/bin", OPENAI_API_KEY: "test-only-not-a-key", OPENAI_BASE_URL: "https://example.invalid" } });
expect(invocation.env).toEqual({ PATH: "/usr/bin" });
expect(invocation.args).toContain("--search");
expect(invocation.args).toContain("read-only");
expect(invocation.args).toContain("skip_host_skill_discovery");
expect(invocation.args[invocation.args.indexOf("--model") + 1]).toBe(NEWS_WRITER_MODEL);
expect(invocation.args).toContain('model_reasoning_effort="medium"');
for (const feature of ["plugins", "apps", "multi_agent", "shell_tool"]) {
expect(invocation.args[invocation.args.indexOf(feature) - 1]).toBe("--disable");
}
const translator = createCodexWriterInvocation({ schemaPath: "schema.json", outputPath: "article.json", purpose: "translation", env: {} });
expect(translator.args).not.toContain("--search");
expect(translator.args).toContain('web_search="disabled"');
expect(translator.args[translator.args.indexOf("--model") + 1]).toBe(NEWS_WRITER_MODEL);
expect(translator.args).toContain('model_reasoning_effort="low"');
});

test("writer records usage without article text or credentials and removes its temporary directory", async () => {
Expand All @@ -31,7 +35,7 @@ test("writer records usage without article text or credentials and removes its t
});
expect(JSON.parse(result)).toEqual({ translation: "fixture" });
const usage = JSON.parse(await readFile(usageFile, "utf8"));
expect(usage).toMatchObject({ purpose: "translation", inputTokens: 100, cachedInputTokens: 30, outputTokens: 20 });
expect(usage).toMatchObject({ purpose: "translation", model: NEWS_WRITER_MODEL, inputTokens: 100, cachedInputTokens: 30, outputTokens: 20 });
expect(JSON.stringify(usage)).not.toContain("private test input");
await expect(readFile(join(temporary, "article.json"))).rejects.toMatchObject({ code: "ENOENT" });
} finally { await rm(directory, { recursive: true, force: true }); }
Expand Down
Loading