ai_synth

Commit Graph

Author	SHA1	Message	Date
oabrivard	eba721266f	feat: article history entry struct + insert/query/cleanup functions Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	d7afd08eaf	feat: enrich article_history with tracing metadata + syntheses.job_id	3 months ago
oabrivard	7cbb2853ce	feat: Autre fill-up to 75% synthesis target with source diversity enforcement Accumulates overflow articles from both classification phases and redistributes them into the Autre category when total articles fall below 75% of the configured max, respecting per-source diversity limits. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	c3e6103ef1	feat: parse_classification_response collects overflow articles Returns a (result, overflow) tuple so callers can access articles that could not fit in any category or Autre. Also adds the SYNTHESIS_MIN_FILL_RATIO constant for the upcoming fill-up logic. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	cea723f7d7	test: update E2E and integration tests with article_history_days setting	3 months ago
oabrivard	65eb6004d2	feat: article history filtering in pipeline — cleanup, Phase 1/2 filter, retry, insert Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	0a87b7ed8f	feat: add normalize_article_url and hash_article_url utilities Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	5a928aa990	feat: add article_history DB module (check, insert, cleanup) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	c271c240a2	feat: add article_history table and article_history_days setting Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	8e06357b47	test: update integration test with LLM scraping settings	3 months ago
oabrivard	8a061c98db	feat: LLM-assisted article extraction with Arc provider, concurrency control, and progress Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	357f06e405	feat: LLM-assisted source link extraction with heuristic fallback Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	e6e8aa1eeb	feat: add LLM prompts and schemas for link and article extraction Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	23f121a58d	feat: ScrapedContent url+head_html fields, Arc<dyn LlmProvider>, 3-tuple scrape returns Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	e483789d1b	feat: add use_llm_for_source_links and use_llm_for_article_extraction settings Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	53ecce84b0	feat: two-phase generation pipeline — personalized sources first, web search fallback Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	51ea032838	feat: add scrape_flat_urls helper and gap-aware search prompt Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	d508b5b4ab	feat: Autre category support in rewrite schema, final sections, URL restore + remove dead code Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	ba7024e280	feat: add classification response parsing with category filling and Autre fallback Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	104b6a0d7b	feat: add classification prompt and schema for article categorization Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	c06b5ba454	feat: add source_scraper module for extracting article links from source pages Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	45e5ee8a7d	fix: rewrite pass schema uses actual scraped item counts, not max setting The rewrite pass shared the search pass schema which enforced minItems/maxItems equal to max_items_per_category. After filter_empty_scraped_articles removed old/failed articles, the scraped data had fewer items than the schema required, causing the LLM to duplicate content to fill the quota. Now build_rewrite_schema counts actual items per category from the scraped data and sets minItems/maxItems accordingly. Also removed dead domain_counts variable. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	13894a8f50	fix: filter empty scraped articles + restore URLs after rewrite + E2E assertions - filter_empty_scraped_articles: removes articles with empty scraped content (too old, soft 404, scrape failure) before the rewrite pass, preventing empty articles in the final synthesis - restore_scraped_urls: already existed, now has unit tests - E2E test: added assertions for no Wikipedia URLs, no empty summaries, and updated settings payload with new fields (max_articles_per_source, source_diversity_window) - 4 new unit tests for filter_empty + restore_scraped_urls Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	a9be1ce435	fix: restore scraped URLs after LLM rewrite pass to prevent hallucination The rewrite pass can replace validated URLs with hallucinated ones (Wikipedia, corporate sites) despite being instructed to preserve them. After the rewrite, restore_scraped_urls() replaces each article's URL with the original scraped URL by matching on position (category + item index). Logs when a URL is restored so hallucination patterns can be monitored. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	8a18b70aff	fix: set max output tokens to 16384 for all LLM providers OpenAI's default output limit (4096 tokens) was too low for structured synthesis output with multiple categories and articles per category, causing truncated JSON. Set 16384 for both OpenAI APIs (Responses + Chat Completions) and Gemini. Anthropic already had 16384. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	55c2b050b3	feat: extract recent domains and pass to search prompt for diversity Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	3f6ad9853c	feat: build_search_prompt accepts recent_domains for source diversity Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	a31915d3ce	feat: add source_diversity_window setting (migration + model + DB + validation tests) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	b558619d10	feat: source diversity limit + URL deduplication in generation pipeline - Add max_articles_per_source setting (default 3, range 1-10) with migration, backend model, DB queries, and frontend number input - Add limit_articles_per_source filter: spreads articles across categories (1 per domain per category first), then fills remaining slots up to the limit - Add dedup_by_url filter: removes duplicate URLs across categories (case-insensitive) - Pipeline order: parse → filter_homepage → dedup_by_url → limit_per_source → scrape - 10 new unit tests covering spread, cap enforcement, dedup, and edge cases Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	6819c7193c	feat: add limit_articles_per_source filter with unit tests Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	c1ee79bcf6	feat: add max_articles_per_source setting (migration + model + DB) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	a3f4c3b42f	fix: always run scrape+rewrite pass to prevent hallucinated URLs The adaptive pipeline skipped the scrape+rewrite pass when the LLM's search results had URLs starting with "http". But LLMs hallucinate plausible URLs (Wikipedia, corporate sites) that pass the http check but aren't actual source articles. The scrape pass catches these by fetching each URL and validating the content exists. Always running the full pipeline ensures URL integrity. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	45c9e71589	fix: enforce max_items_per_category in JSON schema and prompt The LLM was returning only 1 article per category despite the user setting 4. - Added minItems/maxItems to the category array schema (enforced by OpenAI strict mode) - Changed prompt from "au maximum N actualites" to "exactement N actualites" - Schema builder now takes max_items_per_category parameter Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	0b0702de39	fix: strip null bytes from LLM output before saving to PostgreSQL JSONB LLM output occasionally contains \u0000 null bytes (e.g., "annonc\u0000...") which PostgreSQL rejects in JSONB columns. Added sanitize_json_null_bytes() that recursively strips null bytes from all string values before DB insert. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	3fe667591d	fix: LLM providers use own HTTP client with 120s timeout (was sharing scraper's 15s) The scraper client (build_scraper_client) has a 15s timeout appropriate for web scraping, but LLM API calls — especially with web search — take 30-60s. LLM providers now build their own reqwest client with 120s timeout via build_llm_client(). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	6fe75d77e7	feat: add source file:line to WARN and ERROR log lines Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	004f08f385	fix: runtime bugs found during first Docker run + integration tests Bugs fixed: - resolve_model queried non-existent admin_provider_models table (use JSONB query on admin_providers) - key_prefix VARCHAR(10) too short for 11-char prefix (migration to VARCHAR(12)) - API key test schema missing additionalProperties: false (OpenAI strict mode) - CSP missing font-src data: directive (PDF font embedding blocked) - Magic link URL not logged in test mode (can't verify without real email) - Rust 1.85 Docker image too old for dependencies (bumped to 1.88) Tests added to prevent recurrence: - schema_meets_openai_strict_mode_requirements: validates additionalProperties on all objects - key_prefix_full_length_stored_in_db: verifies 11-char prefix survives DB round-trip - generate_pipeline_resolves_model_from_admin_config: exercises full generation pipeline Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	069a4f2022	feat: graceful shutdown and frontend build in Docker - Add SIGTERM/Ctrl+C signal handling with graceful connection draining - Close database pool cleanly on shutdown - Add frontend-builder stage to Dockerfile (node:22-alpine, npm ci + build) - Move Docker build context to project root so both frontend/ and backend/ are accessible - Frontend dist/ copied into container at ./static/ for the backend to serve Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	b961f82f01	refactor: add UserRateLimitEntry constructor and settings_changed method Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	c1f2f1456f	refactor: simplify recent changes — extract helper, named struct, atomic entry, pre-alloc - Extract auth::create_and_send_magic_link() to deduplicate token rollback logic - Replace (i32, i32, RateLimiter) tuple with named UserRateLimitEntry struct - Use DashMap entry API for atomic rate limiter lookup (fixes TOCTOU race) - Pre-allocate scraper body Vec from Content-Length when available Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	54d54f2a06	fix: architect assessment remediation — 6 issues across backend, frontend, and infra - Wire hardened scraper client into runtime (SSRF redirect validation was defined but unused) - Stream scraper body with per-chunk size limit instead of post-download check (DoS/OOM) - Persist user rate-limit overrides across generation jobs via AppState DashMap - Roll back magic-link token on email send failure to prevent quota exhaustion - Fix API error UX: prefer human message over machine error code in frontend - Unwrap GET /syntheses { items } wrapper in frontend API layer (contract mismatch) - Bind Postgres to localhost in docker-compose (was exposed on all interfaces) - Fix CLAUDE.md: runtime queries not compile-time, 10 migrations not 9 Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	ae01bc8e62	security: SSRF redirect validation per hop with custom reqwest policy Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	a4e618feda	test: add unit tests for auth middleware cookie extraction Extract cookie parsing into a standalone `extract_session_token` function and add 5 unit tests covering the valid, missing, multi-cookie, whitespace, and empty-header cases. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	98528f51bd	Fix rate limiter bug, simplify v2 code Bug fix: - Per-generation rate limiter was creating a new instance on every check, making user rate limit overrides non-functional. Fixed by creating the limiter once at pipeline start and reusing for both passes. Simplifications: - Extract spawn_task closure in scrape_articles (deduplicate spawn blocks) - Use idiomatic if let Ok(...) instead of if let Some(..).ok() in scraper - Replace manual loop with iterator chain in export_keys handler - Simplify check_rate_limit to single boolean check - Simplify handleImport settings merge (spread already provides defaults) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	0f66c28c38	v2: empty sections fallback in email template Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	7eb24cfd9a	v2: API key export endpoint (POST, rate-limited) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	3 months ago
oabrivard	191e1c716b	v2: enhanced scraper - title priority chain, broken link detection, noindex Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	9b994e0528	v2: pipeline user model selection, rate limiter, URL filter, original title, null-safe sections - resolve_provider_and_key() now respects user ai_provider preference - Dual model resolution: ai_model for search pass, ai_model_writing for rewrite pass - Per-generation rate limiter with user override support - Homepage URL filter removes domain-only URLs after search pass - ScrapedNewsItem gains original_title field populated from page <title> - SynthesisResponse::try_from handles null sections gracefully (returns empty vec) - Search prompt warns LLM against returning homepage URLs - Rewrite prompt instructs LLM to use originalTitle with language preservation rules Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	ed6b41fe52	v2: add settings migration, model expansion, DB queries (provider, models, rate limits) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago
oabrivard	04819aa926	Simplify code: deduplicate patterns, fix captcha field name bug Bug fix: - Fix frontend sending captcha_token instead of turnstile_token in login/register requests (would cause 422 errors on auth) Backend simplifications: - Deduplicate VALID_PROVIDERS constant (provider.rs is now the single source) - Extract validate_display_name/validate_models helpers in provider model - Add From<UserSettings> for SettingsResponse, From<User> for AdminUserResponse - Consolidate Resend API call pattern into shared send_via_resend() - Extract do_bulk_import() for sources bulk/CSV import - Use idiomatic range.contains() for rate limit validation Frontend simplifications: - Consolidate file download logic (exportCsv reuses shared fetchFile/triggerDownload) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	3 months ago

1 2

59 Commits (eba721266f2bf424d768e4549f68ff20436991c3)