Compare commits
12 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| bda4a753c9 | |||
| de91485854 | |||
| 2fa50a8600 | |||
| df7bf98ef1 | |||
| 159080f0c8 | |||
| e4d2073eeb | |||
| 6dc847a151 | |||
| e1a14ec3fc | |||
| c5dcbdb997 | |||
| 4bc86b9494 | |||
| 7c5330a528 | |||
| 39256c0f9b |
@@ -17,6 +17,10 @@ joomla-php83/
|
|||||||
backups/*
|
backups/*
|
||||||
!backups/README.md
|
!backups/README.md
|
||||||
|
|
||||||
|
# Backups pesados (GBs) bajo docs/ — mismo motivo que backups/*, ruta distinta
|
||||||
|
docs/backups/*
|
||||||
|
!docs/backups/README.md
|
||||||
|
|
||||||
# Capturas de pantalla
|
# Capturas de pantalla
|
||||||
capturas/
|
capturas/
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,219 @@
|
|||||||
|
# GA4 API setup for feadulta
|
||||||
|
|
||||||
|
This document describes the simplest practical path for querying Google Analytics 4 from this repo.
|
||||||
|
|
||||||
|
> **Where the code lives vs where it runs (2026-07-31).** This script used to live only in the
|
||||||
|
> separate `feadulta-git` checkout, which points at the *archived* Gitea and never made it into
|
||||||
|
> this repo. The canonical copy is now here, in `rafa/feadulta` on `gitea.feadulta.com`.
|
||||||
|
> The **runtime environment stays in `/mnt/c/Users/Chia/feadulta-git`**: `.venv/` and, above all,
|
||||||
|
> `.secrets/` (OAuth client + cached token) are gitignored and were never versioned anywhere.
|
||||||
|
> That is why the commands below still use absolute paths into `feadulta-git` — the paths are
|
||||||
|
> correct, the code is just no longer only there.
|
||||||
|
|
||||||
|
## Current known identifier
|
||||||
|
|
||||||
|
The site is tagged with GA4 measurement ID:
|
||||||
|
|
||||||
|
- `G-6RT9ZRS4LW`
|
||||||
|
|
||||||
|
Important:
|
||||||
|
|
||||||
|
- the GA4 **measurement ID** (`G-...`) is **not** the same as the GA4 **property ID**
|
||||||
|
- the Data API `runReport` endpoint needs the **property ID**
|
||||||
|
- the script added in this repo can resolve the property automatically if the authenticated Google user has access to the property
|
||||||
|
|
||||||
|
Official references:
|
||||||
|
|
||||||
|
- Data API `runReport`: https://developers.google.com/analytics/devguides/reporting/data/v1/rest/v1beta/properties/runReport
|
||||||
|
- Admin API overview: https://developers.google.com/analytics/devguides/config/admin/v1
|
||||||
|
- Where to find the measurement ID in GA4: https://support.google.com/analytics/answer/9304153
|
||||||
|
|
||||||
|
## Recommended auth model
|
||||||
|
|
||||||
|
Use **OAuth desktop app credentials** for a Google user that already has access to the GA4 property.
|
||||||
|
|
||||||
|
Why this is the easiest first step:
|
||||||
|
|
||||||
|
- no need to create a service account and grant property access separately
|
||||||
|
- no need to know the property ID upfront
|
||||||
|
- the script can authenticate as you and search the accessible properties for the matching `G-...`
|
||||||
|
|
||||||
|
## One-time Google Cloud setup
|
||||||
|
|
||||||
|
1. Open Google Cloud Console.
|
||||||
|
2. Create or reuse a project.
|
||||||
|
3. Enable:
|
||||||
|
- Google Analytics Data API
|
||||||
|
- Google Analytics Admin API
|
||||||
|
4. Create an OAuth client of type `Desktop app`.
|
||||||
|
5. Download the client secrets JSON file.
|
||||||
|
|
||||||
|
Suggested local path:
|
||||||
|
|
||||||
|
- `/mnt/c/Users/Chia/feadulta-git/.secrets/ga4-oauth-client.json`
|
||||||
|
|
||||||
|
Do not commit it.
|
||||||
|
|
||||||
|
## Local Python environment
|
||||||
|
|
||||||
|
This repo is set up to use a local virtualenv so the host Python installation does not need to be modified.
|
||||||
|
|
||||||
|
Create it once:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 -m venv /mnt/c/Users/Chia/feadulta-git/.venv
|
||||||
|
```
|
||||||
|
|
||||||
|
Install the required packages inside that environment:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python -m pip install google-auth google-auth-oauthlib requests
|
||||||
|
```
|
||||||
|
|
||||||
|
## Environment variables
|
||||||
|
|
||||||
|
You can configure the script with environment variables:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export GA4_CLIENT_SECRETS_PATH=/mnt/c/Users/Chia/feadulta-git/.secrets/ga4-oauth-client.json
|
||||||
|
export GA4_TOKEN_PATH=/mnt/c/Users/Chia/feadulta-git/.secrets/ga4-token.json
|
||||||
|
export GA4_MEASUREMENT_ID=G-6RT9ZRS4LW
|
||||||
|
export GA4_PROPERTY_ID=508378818
|
||||||
|
```
|
||||||
|
|
||||||
|
If `GA4_PROPERTY_ID` is omitted, the script can try to resolve it from `GA4_MEASUREMENT_ID`.
|
||||||
|
|
||||||
|
## First run
|
||||||
|
|
||||||
|
Authenticate and resolve the property:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --measurement-id G-6RT9ZRS4LW resolve-property
|
||||||
|
```
|
||||||
|
|
||||||
|
If the local environment cannot open a browser directly, use manual mode:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --measurement-id G-6RT9ZRS4LW --no-browser resolve-property
|
||||||
|
```
|
||||||
|
|
||||||
|
This prints a Google authorization URL. Open it in the browser, sign in with a Google user that has access to the GA4 property, and complete the redirect back to the `localhost` callback URL shown in the command output.
|
||||||
|
|
||||||
|
On successful first run, the script stores a reusable token locally at:
|
||||||
|
|
||||||
|
- `/mnt/c/Users/Chia/feadulta-git/.secrets/ga4-token.json`
|
||||||
|
|
||||||
|
Current known resolved property:
|
||||||
|
|
||||||
|
- measurement ID: `G-6RT9ZRS4LW`
|
||||||
|
- property ID: `508378818`
|
||||||
|
- property name: `https://feadulta.com`
|
||||||
|
- account name: `Portal feadulta.com`
|
||||||
|
- stream name: `https://www.feadulta.com/`
|
||||||
|
|
||||||
|
## Example reports
|
||||||
|
|
||||||
|
Traffic overview:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset traffic --days 28
|
||||||
|
```
|
||||||
|
|
||||||
|
Top content:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset content --days 28 --limit 25
|
||||||
|
```
|
||||||
|
|
||||||
|
Landing pages:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset landing-pages --days 28 --limit 25
|
||||||
|
```
|
||||||
|
|
||||||
|
Traffic by source / medium:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset source-medium --days 28 --limit 25
|
||||||
|
```
|
||||||
|
|
||||||
|
Device mix:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset device --days 28 --limit 25
|
||||||
|
```
|
||||||
|
|
||||||
|
Export to CSV:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset content --days 28 --csv /tmp/ga4-content.csv
|
||||||
|
```
|
||||||
|
|
||||||
|
## Splitting the live site from the static archive (`--host`)
|
||||||
|
|
||||||
|
This single property (`G-6RT9ZRS4LW`) collects several hostnames at once: the live
|
||||||
|
WordPress (`www.feadulta.com`), the frozen Joomla archive (`antiguo.feadulta.com`,
|
||||||
|
which carries the same GA tag inside its captured HTML), plus leftovers like
|
||||||
|
`wp-nuevo.feadulta.com`. **Any report without a host filter mixes them and means
|
||||||
|
nothing.**
|
||||||
|
|
||||||
|
Which hostnames are actually reporting:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset hosts --days 28
|
||||||
|
```
|
||||||
|
|
||||||
|
Only the live site:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset content --host www.feadulta.com --days 28 --limit 25
|
||||||
|
```
|
||||||
|
|
||||||
|
Only the archive:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/mnt/c/Users/Chia/feadulta-git/.venv/bin/python scripts/ga4_report.py --property-id 508378818 report --preset content --host antiguo.feadulta.com --days 28 --limit 25
|
||||||
|
```
|
||||||
|
|
||||||
|
`--host` takes a comma-separated list (exact match, case-insensitive) and combines
|
||||||
|
with `--page-path-regex` as an AND group. `--host-not` negates it.
|
||||||
|
|
||||||
|
## Practical future access
|
||||||
|
|
||||||
|
For future use, the shortest path is:
|
||||||
|
|
||||||
|
1. Confirm these files still exist locally:
|
||||||
|
- `/mnt/c/Users/Chia/feadulta-git/.secrets/ga4-oauth-client.json`
|
||||||
|
- `/mnt/c/Users/Chia/feadulta-git/.secrets/ga4-token.json`
|
||||||
|
- `/mnt/c/Users/Chia/feadulta-git/.venv/`
|
||||||
|
2. Run reports directly with `--property-id 508378818`.
|
||||||
|
3. Only rerun `resolve-property` if the token was deleted or the Google access changed.
|
||||||
|
4. If the token expires, the script should refresh it automatically when possible.
|
||||||
|
|
||||||
|
## About WordPress logs
|
||||||
|
|
||||||
|
If the question is “what content is being seen?”, GA4 is usually the better first tool because it gives:
|
||||||
|
|
||||||
|
- page-level views
|
||||||
|
- landing pages
|
||||||
|
- traffic sources
|
||||||
|
- device mix
|
||||||
|
- trends over time
|
||||||
|
|
||||||
|
WordPress itself does **not** log page views by default in a way that is useful for editorial analysis.
|
||||||
|
|
||||||
|
If GA4 turns out to be incomplete or unreliable, the next fallback is usually:
|
||||||
|
|
||||||
|
1. web server access logs
|
||||||
|
2. reverse proxy logs
|
||||||
|
3. plugin-specific event logging if the site has a dedicated analytics plugin
|
||||||
|
|
||||||
|
In this repo, there is no obvious WordPress analytics plugin configuration under `wordpress/wp-content/mu-plugins/`, so GA4 or server logs are the most likely useful sources.
|
||||||
|
|
||||||
|
## Useful questions this script should answer
|
||||||
|
|
||||||
|
- Which pages got the most views in the last 28 days?
|
||||||
|
- Which landing pages attract the most traffic?
|
||||||
|
- Which sources or source/medium pairs bring traffic?
|
||||||
|
- Is mobile traffic increasing or decreasing?
|
||||||
|
- Did traffic fall because fewer users arrived, or because fewer pages were viewed per session?
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
# Release dry-run — Enrique Martínez Lozano / carta «Hacia el corazón»
|
||||||
|
|
||||||
|
**Estado:** preparado localmente; **no ejecutado en producción**.
|
||||||
|
|
||||||
|
## Validación local completada
|
||||||
|
|
||||||
|
- Comentario ES: `#55041` — `El tesoro está ya en nosotros`
|
||||||
|
- Categorías: `Comentarios al evangelio` + `Feadulta`.
|
||||||
|
- Audio: `tts/55041.mp3`, 2,221,101 bytes, voz `NicoFeadulta2026`.
|
||||||
|
- QA Haiku: aprobado tras correcciones objetivas EN/FR/IT.
|
||||||
|
- Traducciones del comentario: EN `#55047`, FR `#55048`, IT `#55049`, PT `#55050`.
|
||||||
|
- Carta local ES: `#55042` — `Hacia el corazón` (slug técnico de preview conservado).
|
||||||
|
- Traducciones: EN `#55051`, FR `#55052`, IT `#55053`, PT `#55054`.
|
||||||
|
- QA Haiku: contenido aprobado tras correcciones; 24 enlaces con destino traducido se repuntaron. Los 12 enlaces externos que quedan son fallback ES: no existe post destino traducido en local (verificado por `post_title` exacto y Polylang).
|
||||||
|
- Render comprobado en Tailscale:
|
||||||
|
- Comentario ES: título correcto, categorías correctas y reproductor HTML5 visible.
|
||||||
|
- Carta FR: título correcto; Enrique aparece tras Fray Marcos y antes de Pagola.
|
||||||
|
|
||||||
|
## Evidencia de producción (sólo lectura server-side)
|
||||||
|
|
||||||
|
- Carta real: `#54914`, título `Hacia el corazón`, estado `publish`.
|
||||||
|
- Hash actual de contenido prod: `5e7eb1b880bfc6311edb78ac8b6e81639f5b4502282a1b182c610b1d01308b69`.
|
||||||
|
- Hash candidato local: `6fdb4929fcb5b49b7f2dc1deb7c6f5be70983af866d44c5d5c27d2866c09418c`.
|
||||||
|
- La carta candidata lleva `fea_phase_a_source_prod_id=54914`; por tanto debe **actualizar #54914**, no crearse una carta nueva.
|
||||||
|
|
||||||
|
## Dry-runs que pasaron
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Crea solamente el grupo nuevo del comentario (ES + EN/FR/IT/PT): 5 posts draft
|
||||||
|
python3 scripts/sync_translations_to_prod.py --ids 55041 --dry-run
|
||||||
|
|
||||||
|
# Sube y asocia solamente el MP3 de Enrique
|
||||||
|
python3 scripts/sync_audio_to_prod.py --ids 55041 --dry-run
|
||||||
|
```
|
||||||
|
|
||||||
|
Ambos devolvieron `ok/error=1/0` (audio) y un grupo Polylang de cinco posts (comentario).
|
||||||
|
|
||||||
|
## Guardarraíl crítico
|
||||||
|
|
||||||
|
**NO ejecutar:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 scripts/sync_translations_to_prod.py --ids 55041,55042
|
||||||
|
```
|
||||||
|
|
||||||
|
El dry-run demostró que clonaría `#55042` como una segunda carta ES. No actualiza la carta real `#54914`.
|
||||||
|
|
||||||
|
## Release que debe ejecutar Rafa (producción)
|
||||||
|
|
||||||
|
1. Backup de producción y snapshot server-side de `#54914`.
|
||||||
|
2. Crear el grupo nuevo del comentario desde `#55041` (ES + EN/FR/IT/PT), inicialmente en `draft`; anotar los IDs reales de producción.
|
||||||
|
3. Subir `tts/55041.mp3` y asociarlo al ID ES real creado en producción.
|
||||||
|
4. Actualizar **solamente** título/contenido de la carta real `#54914` con el candidato local `#55042`; mantener título `Hacia el corazón`, autor, fecha y estado existentes.
|
||||||
|
5. Crear EN/FR/IT/PT de la carta contra el origen real `#54914` y enlazarlas en Polylang. No clonar otra ES.
|
||||||
|
6. Releer por servidor los IDs creados y `#54914`; comprobar:
|
||||||
|
- Enrique sólo está en `Evangelio y comentarios al Evangelio`, tras Fray Marcos y antes de Pagola.
|
||||||
|
- No aparece en `Artículos seleccionados`.
|
||||||
|
- Cada carta enlaza a Enrique en su mismo idioma.
|
||||||
|
- El reproductor del ES carga `tts/55041.mp3`.
|
||||||
|
7. Sólo tras esas lecturas, promover los drafts de producción a `publish`.
|
||||||
|
|
||||||
|
## Rollback
|
||||||
|
|
||||||
|
- Restaurar el backup de producción previo.
|
||||||
|
- Alternativamente, restaurar el contenido/hash previo de `#54914` y despublicar el nuevo grupo + audio de Enrique.
|
||||||
|
|
||||||
|
## Backups locales relacionados
|
||||||
|
|
||||||
|
- `docs/backups/enrique-translations-20260725T004921Z/`
|
||||||
|
- `docs/backups/enrique-carta-fr-final-20260725T012649Z/`
|
||||||
|
- `docs/backups/enrique-pre-tts-20260725T013021Z/`
|
||||||
@@ -0,0 +1,50 @@
|
|||||||
|
<?php
|
||||||
|
/**
|
||||||
|
* Issue #62 — Reapunta foto_perfil de cada autor a su avatar nuevo.
|
||||||
|
* Crea un attachment por autor y guarda el foto_perfil antiguo en _foto_perfil_pre62 (revertible).
|
||||||
|
* Uso: wp eval-file import_avatars_62.php [--apply]
|
||||||
|
* sin --apply => dry run (no escribe nada).
|
||||||
|
*/
|
||||||
|
require_once ABSPATH . 'wp-admin/includes/image.php';
|
||||||
|
|
||||||
|
$apply = getenv('APPLY') === '1';
|
||||||
|
$mfile = getenv('MANIFEST') ?: 'manifest_62.json';
|
||||||
|
$updir = wp_get_upload_dir();
|
||||||
|
$manifest = json_decode(file_get_contents($updir['basedir'] . '/avatares/' . $mfile), true);
|
||||||
|
|
||||||
|
$done = $skip = $err = 0;
|
||||||
|
foreach ($manifest as $row) {
|
||||||
|
list($uid, $name, $orig, $kind, $old_attach) = $row;
|
||||||
|
if (!in_array($kind, ['initials','trim','logo','qs'], true)) { $skip++; continue; }
|
||||||
|
|
||||||
|
$rel = "avatares/autores/autor-{$uid}.png";
|
||||||
|
$abs = $updir['basedir'] . '/' . $rel;
|
||||||
|
if (!file_exists($abs)) { echo "MISSING file uid=$uid\n"; $err++; continue; }
|
||||||
|
|
||||||
|
// si foto_perfil ya apunta a este fichero, solo se ha sobrescrito el PNG: no crear attachment nuevo
|
||||||
|
$cur = (int) get_user_meta($uid, 'foto_perfil', true);
|
||||||
|
if ($cur && get_post_meta($cur, '_wp_attached_file', true) === $rel) {
|
||||||
|
if ($apply) wp_update_attachment_metadata($cur, wp_generate_attachment_metadata($cur, $abs));
|
||||||
|
$skip++; continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!$apply) { $done++; continue; }
|
||||||
|
|
||||||
|
// guardar foto_perfil antiguo una sola vez (idempotente)
|
||||||
|
if (get_user_meta($uid, '_foto_perfil_pre62', true) === '') {
|
||||||
|
update_user_meta($uid, '_foto_perfil_pre62', $old_attach);
|
||||||
|
}
|
||||||
|
|
||||||
|
$attach = [
|
||||||
|
'post_mime_type' => 'image/png',
|
||||||
|
'post_title' => "Avatar {$name}",
|
||||||
|
'post_status' => 'inherit',
|
||||||
|
'guid' => $updir['baseurl'] . '/' . $rel,
|
||||||
|
];
|
||||||
|
$aid = wp_insert_attachment($attach, $abs, 0, true);
|
||||||
|
if (is_wp_error($aid)) { echo "ERR insert uid=$uid: ".$aid->get_error_message()."\n"; $err++; continue; }
|
||||||
|
wp_update_attachment_metadata($aid, wp_generate_attachment_metadata($aid, $abs));
|
||||||
|
update_user_meta($uid, 'foto_perfil', $aid);
|
||||||
|
$done++;
|
||||||
|
}
|
||||||
|
echo ($apply ? "APLICADO" : "DRY-RUN") . ": procesados=$done omitidos=$skip errores=$err\n";
|
||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
File diff suppressed because one or more lines are too long
@@ -13,6 +13,21 @@ server {
|
|||||||
# application/octet-stream y el navegador se los descarga en vez de mostrarlos.
|
# application/octet-stream y el navegador se los descarga en vez de mostrarlos.
|
||||||
default_type text/html;
|
default_type text/html;
|
||||||
|
|
||||||
|
# URLs históricas de catálogo: el catálogo vigente vive en Ediciones Fe Adulta.
|
||||||
|
# 302 primero: evita cachear un destino externo de forma irreversible durante el soak.
|
||||||
|
location ~ ^/(?:es/)?catalogo-de-libros-feadulta(?:/|\.html)?$ {
|
||||||
|
return 302 https://edicionesfeadulta.com/;
|
||||||
|
}
|
||||||
|
|
||||||
|
# La vista imprimible Joomla de esta página EFFA no existe en el archivo estático.
|
||||||
|
# Conserva la página normal del archivo; sólo la variante ?print=1 va al contenido WP vigente.
|
||||||
|
location = /es/effa/95-secc5cat/3304-seccion5col00.html {
|
||||||
|
if ($arg_print = 1) {
|
||||||
|
return 302 https://www.feadulta.com/seccion5col00-2/;
|
||||||
|
}
|
||||||
|
try_files $uri $uri.html $uri/index.html =404;
|
||||||
|
}
|
||||||
|
|
||||||
location / {
|
location / {
|
||||||
try_files $uri $uri.html $uri/index.html =404;
|
try_files $uri $uri.html $uri/index.html =404;
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,10 @@
|
|||||||
|
portfolio-tracker
|
||||||
|
yt-api-bridge
|
||||||
|
wordpress-web
|
||||||
|
joomla-web-php83
|
||||||
|
jellyfin
|
||||||
|
triptyk-local-wordpress-1
|
||||||
|
triptyk-local-db-1
|
||||||
|
yt-summaries
|
||||||
|
joomla-web
|
||||||
|
wordpress-mysql
|
||||||
Executable
+37
@@ -0,0 +1,37 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Retira exclusivamente el esquema vacío fewp1 de CDMON.
|
||||||
|
# Por defecto solo hace dry-run. La eliminación real exige --apply y el token exacto.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
mode="dry-run"
|
||||||
|
if [[ "${1:-}" == "--apply" ]]; then
|
||||||
|
mode="apply"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ "$mode" == "apply" && "${CONFIRM_DROP_FEWP1:-}" != "DROP_FEWP1" ]]; then
|
||||||
|
echo "Refusing apply: export CONFIRM_DROP_FEWP1=DROP_FEWP1 first." >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
set -a
|
||||||
|
source "$HOME/.hermes/profiles/feadulta/.env"
|
||||||
|
set +a
|
||||||
|
|
||||||
|
php='global $wpdb;
|
||||||
|
$count=(int)$wpdb->get_var("SELECT COUNT(*) FROM information_schema.tables WHERE table_schema=\"fewp1\"");
|
||||||
|
if ($count !== 0) { fwrite(STDERR, "ABORT: fewp1 has ".$count." tables; it is not empty.\n"); exit(3); }
|
||||||
|
echo "PASS: fewp1 exists and has 0 tables.\n";'
|
||||||
|
|
||||||
|
if [[ "$mode" == "apply" ]]; then
|
||||||
|
php+=' $ok=$wpdb->query("DROP DATABASE `fewp1`");
|
||||||
|
if ($ok === false) { fwrite(STDERR, "ABORT: DROP DATABASE failed.\n"); exit(4); }
|
||||||
|
$remaining=(int)$wpdb->get_var("SELECT COUNT(*) FROM information_schema.schemata WHERE schema_name=\"fewp1\"");
|
||||||
|
if ($remaining !== 0) { fwrite(STDERR, "ABORT: fewp1 remains after DROP.\n"); exit(5); }
|
||||||
|
echo "APPLIED: fewp1 removed; verified absent.\n";'
|
||||||
|
else
|
||||||
|
php+=' echo "DRY-RUN: would execute DROP DATABASE `fewp1`; no production change made.\n";'
|
||||||
|
fi
|
||||||
|
|
||||||
|
printf -v quoted_php '%q' "$php"
|
||||||
|
sshpass -p "$FEA_PROD_SSH_PASS" ssh -o StrictHostKeyChecking=accept-new "$FEA_PROD_SSH_HOST" \
|
||||||
|
"cd /web && wp eval $quoted_php"
|
||||||
+58
-1
@@ -3,11 +3,16 @@
|
|||||||
* IO mínimo de posts WP para el reprocesador EN.
|
* IO mínimo de posts WP para el reprocesador EN.
|
||||||
* get <id> -> escribe /tmp/fea_es.json {title, content, status}
|
* get <id> -> escribe /tmp/fea_es.json {title, content, status}
|
||||||
* update <id> <titlef> <bodyf> -> actualiza post_title/post_content desde ficheros
|
* update <id> <titlef> <bodyf> -> actualiza post_title/post_content desde ficheros
|
||||||
|
* listpending <autor> ... -> cola de backlog TTS pendiente (ver abajo)
|
||||||
* Carga wp-load; portable (local docker o prod via FEA_WP_LOAD).
|
* Carga wp-load; portable (local docker o prod via FEA_WP_LOAD).
|
||||||
*/
|
*/
|
||||||
$WP = getenv('FEA_WP_LOAD') ?: '/var/www/html/wp-load.php';
|
$WP = getenv('FEA_WP_LOAD') ?: '/var/www/html/wp-load.php';
|
||||||
require $WP;
|
require $WP;
|
||||||
|
|
||||||
|
// Por debajo de esto el post_content no da para locutar (prefiltro barato en SQL;
|
||||||
|
// tts_produce.py vuelve a medir el texto ya extraído y marca fea_audio_skip).
|
||||||
|
const FEA_TTS_MIN_CONTENT = 400;
|
||||||
|
|
||||||
$action = $argv[1] ?? '';
|
$action = $argv[1] ?? '';
|
||||||
|
|
||||||
if ($action === 'get') {
|
if ($action === 'get') {
|
||||||
@@ -19,6 +24,8 @@ if ($action === 'get') {
|
|||||||
'title' => $p->post_title,
|
'title' => $p->post_title,
|
||||||
'content' => $p->post_content,
|
'content' => $p->post_content,
|
||||||
'status' => $p->post_status,
|
'status' => $p->post_status,
|
||||||
|
'post_type' => $p->post_type,
|
||||||
|
'post_name' => $p->post_name,
|
||||||
'author' => (int)$p->post_author,
|
'author' => (int)$p->post_author,
|
||||||
], JSON_UNESCAPED_UNICODE));
|
], JSON_UNESCAPED_UNICODE));
|
||||||
exit(0);
|
exit(0);
|
||||||
@@ -70,5 +77,55 @@ if ($action === 'unsetaudio') { // unsetaudio <id> (rollback: despublica el au
|
|||||||
exit(0);
|
exit(0);
|
||||||
}
|
}
|
||||||
|
|
||||||
fwrite(STDERR, "uso: get|update|getmeta|setaudio|setflag|unsetaudio\n");
|
if ($action === 'listpending') { // listpending <autor> <desde> <hasta> <limite> [voz_esperada]
|
||||||
|
// Cola del backlog de TTS: posts ES publicados de un autor que todavía no tienen
|
||||||
|
// audio. La consulta ES la idempotencia — no hay fichero de estado que mantener:
|
||||||
|
// lo ya locutado deja de salir solo. Si se pasa la voz clonada del autor, también
|
||||||
|
// salen los que se locutaron en su día con otra voz (p. ej. Nico), para rehacerlos.
|
||||||
|
$autor = (int)($argv[2] ?? 0);
|
||||||
|
$desde = (int)($argv[3] ?? 0);
|
||||||
|
$hasta = (int)($argv[4] ?? 9999);
|
||||||
|
$limite = (int)($argv[5] ?? 50);
|
||||||
|
$voz = (string)($argv[6] ?? '');
|
||||||
|
if (!$autor) {
|
||||||
|
fwrite(STDERR, "uso: listpending <autor> <desde> <hasta> <limite> [voz_esperada]\n");
|
||||||
|
exit(2);
|
||||||
|
}
|
||||||
|
if ($limite <= 0) { $limite = 100000; }
|
||||||
|
|
||||||
|
global $wpdb;
|
||||||
|
$es = (int)$wpdb->get_var("
|
||||||
|
SELECT tt.term_taxonomy_id FROM {$wpdb->term_taxonomy} tt
|
||||||
|
JOIN {$wpdb->terms} t ON t.term_id = tt.term_id
|
||||||
|
WHERE tt.taxonomy = 'language' AND t.slug = 'es' LIMIT 1");
|
||||||
|
if (!$es) { fwrite(STDERR, "no encuentro el idioma 'es' de polylang\n"); exit(1); }
|
||||||
|
|
||||||
|
$ids = $wpdb->get_col($wpdb->prepare("
|
||||||
|
SELECT p.ID
|
||||||
|
FROM {$wpdb->posts} p
|
||||||
|
JOIN {$wpdb->term_relationships} tr
|
||||||
|
ON tr.object_id = p.ID AND tr.term_taxonomy_id = %d
|
||||||
|
LEFT JOIN {$wpdb->postmeta} done
|
||||||
|
ON done.post_id = p.ID AND done.meta_key = 'fea_audio_done'
|
||||||
|
LEFT JOIN {$wpdb->postmeta} voz
|
||||||
|
ON voz.post_id = p.ID AND voz.meta_key = 'fea_audio_voice'
|
||||||
|
LEFT JOIN {$wpdb->postmeta} skip
|
||||||
|
ON skip.post_id = p.ID AND skip.meta_key = 'fea_audio_skip'
|
||||||
|
WHERE p.post_author = %d
|
||||||
|
AND p.post_type = 'post'
|
||||||
|
AND p.post_status = 'publish'
|
||||||
|
AND YEAR(p.post_date) BETWEEN %d AND %d
|
||||||
|
AND CHAR_LENGTH(p.post_content) >= %d
|
||||||
|
AND (skip.meta_value IS NULL OR skip.meta_value <> '1')
|
||||||
|
AND (done.meta_value IS NULL
|
||||||
|
OR done.meta_value <> '1'
|
||||||
|
OR (%s <> '' AND COALESCE(voz.meta_value, '') <> %s))
|
||||||
|
ORDER BY p.post_date DESC
|
||||||
|
LIMIT %d", $es, $autor, $desde, $hasta, FEA_TTS_MIN_CONTENT, $voz, $voz, $limite));
|
||||||
|
|
||||||
|
foreach ($ids as $id) { echo (int)$id . "\n"; }
|
||||||
|
exit(0);
|
||||||
|
}
|
||||||
|
|
||||||
|
fwrite(STDERR, "uso: get|update|getmeta|setaudio|setflag|unsetaudio|listpending\n");
|
||||||
exit(2);
|
exit(2);
|
||||||
|
|||||||
@@ -262,6 +262,47 @@ switch ($cmd) {
|
|||||||
break;
|
break;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
case 'clone_new': {
|
||||||
|
// Clona en un ID local NUEVO. Se usa al importar desde prod cuando el
|
||||||
|
// mismo ID ya está ocupado por contenido local distinto.
|
||||||
|
$lang = (string) ($argv[2] ?? '');
|
||||||
|
$status = (string) ($argv[3] ?? '');
|
||||||
|
if ($lang === '') {
|
||||||
|
fwrite(STDERR, "uso: clone_new <lang> <status>\n"); exit(6);
|
||||||
|
}
|
||||||
|
$payload = json_decode(file_get_contents('php://stdin'), true);
|
||||||
|
if (!is_array($payload) || empty($payload['title'])) {
|
||||||
|
fwrite(STDERR, "payload inválido por stdin\n"); exit(4);
|
||||||
|
}
|
||||||
|
$postarr = [
|
||||||
|
'post_title' => wp_slash($payload['title']),
|
||||||
|
'post_content' => wp_slash($payload['content'] ?? ''),
|
||||||
|
'post_excerpt' => wp_slash($payload['excerpt'] ?? ''),
|
||||||
|
'post_status' => $status ?: ($payload['status'] ?? 'draft'),
|
||||||
|
'post_type' => $payload['type'] ?? 'post',
|
||||||
|
'post_author' => (int) ($payload['author'] ?? 1),
|
||||||
|
'post_date' => $payload['date'] ?? current_time('mysql'),
|
||||||
|
'post_date_gmt'=> $payload['date_gmt'] ?? current_time('mysql', true),
|
||||||
|
'post_name' => $payload['slug'] ?? '',
|
||||||
|
'to_ping' => '',
|
||||||
|
'pinged' => '',
|
||||||
|
];
|
||||||
|
$new_id = wp_insert_post($postarr, true);
|
||||||
|
if (is_wp_error($new_id)) { fwrite(STDERR, $new_id->get_error_message() . "\n"); exit(5); }
|
||||||
|
pll_set_post_language($new_id, $lang);
|
||||||
|
$cats = [];
|
||||||
|
foreach ((array) ($payload['cat_slugs'] ?? []) as $slug) {
|
||||||
|
$term = get_term_by('slug', (string) $slug, 'category');
|
||||||
|
if ($term && !is_wp_error($term)) $cats[] = (int) $term->term_id;
|
||||||
|
}
|
||||||
|
if (!$cats) $cats = array_values(array_unique(array_map('intval', (array) ($payload['cats'] ?? []))));
|
||||||
|
wp_set_post_categories($new_id, $cats);
|
||||||
|
set_meta_payload($new_id, normalize_meta_input($payload));
|
||||||
|
clean_post_cache($new_id);
|
||||||
|
echo $new_id;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
case 'carta_sections': {
|
case 'carta_sections': {
|
||||||
// Descubre el cluster de una carta parseando sus propios enlaces internos
|
// Descubre el cluster de una carta parseando sus propios enlaces internos
|
||||||
// (misma lógica que pinta la portada, fea-carta-portada.php). No depende
|
// (misma lógica que pinta la portada, fea-carta-portada.php). No depende
|
||||||
|
|||||||
Executable
+257
@@ -0,0 +1,257 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Reporte diario del backlog de TTS de Fe Adulta (issue rafa/feadulta#188).
|
||||||
|
|
||||||
|
SOLO LECTURA: no genera audio ni toca la BD. Cuenta lo hecho en las últimas 24 h,
|
||||||
|
lo que queda por autor, la cuota de MiniMax y las ventanas que se saltaron.
|
||||||
|
|
||||||
|
Entregado por Hermes en modo no-agent (el stdout va directo a Rafa).
|
||||||
|
Silencio deliberado si no hay nada que contar y todo está en orden.
|
||||||
|
|
||||||
|
Hermes solo ejecuta scripts que resuelvan DENTRO de ~/.hermes/scripts, y resuelve
|
||||||
|
los symlinks antes de comprobarlo: un enlace a este fichero se bloquea. Por eso
|
||||||
|
~/.hermes/scripts/fea_tts_backlog_report.py es un wrapper que lo llama por
|
||||||
|
subproceso (mismo patrón que feadulta_ga4_daily.py). Este de aquí es el único
|
||||||
|
sitio donde se edita la lógica.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
from datetime import datetime, timedelta
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
REPO = Path("/home/rafa/joomla-migration")
|
||||||
|
TTS_DIR = REPO / "wordpress/wp-content/uploads/tts"
|
||||||
|
LOG_DIR = Path("/tmp/fea-tts-backlog")
|
||||||
|
QUOTA = Path("/home/rafa/ytsummaries/scripts/quota.py")
|
||||||
|
CONTAINER = "wordpress-web"
|
||||||
|
CRON = "/home/rafa/joomla-migration/scripts/tts_backlog_cron.sh"
|
||||||
|
|
||||||
|
# WP user_id -> (nombre, voz clonada). Mismo mapping que AUTHOR_VOICES en
|
||||||
|
# scripts/minimax_tts.py; si se clona una voz nueva, añadirla en los dos sitios.
|
||||||
|
AUTORES = {
|
||||||
|
382: ("Fray Marcos", "FrayMarcosFeadulta2026"),
|
||||||
|
383: ("Pagola", "PagolaFeadulta2026"),
|
||||||
|
774: ("Sicre", "SicreFeadulta2026"),
|
||||||
|
386: ("Arregi", "ArregiFeadulta2026"),
|
||||||
|
}
|
||||||
|
|
||||||
|
# Autor y rango que está locutando el cron ahora mismo. Debe seguir el crontab:
|
||||||
|
# solo afecta a qué línea se marca como «en curso» en el informe diario.
|
||||||
|
ACTIVO = (774, 2000, 2026)
|
||||||
|
|
||||||
|
|
||||||
|
def php(*args: str) -> str:
|
||||||
|
"""Llama al helper WP sin convertir un fallo de lectura en una cola vacía."""
|
||||||
|
r = subprocess.run(
|
||||||
|
["docker", "exec", CONTAINER, "php", "/tmp/fea_post_io.php", *args],
|
||||||
|
capture_output=True, text=True,
|
||||||
|
)
|
||||||
|
if r.returncode != 0:
|
||||||
|
detail = (r.stderr or r.stdout).strip().replace("\n", " ")[:300]
|
||||||
|
raise RuntimeError(f"consulta WP {' '.join(args)} falló (rc={r.returncode}): {detail}")
|
||||||
|
return r.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def pendientes(autor: int, desde: int, hasta: int, voz: str) -> int:
|
||||||
|
salida = php("listpending", str(autor), str(desde), str(hasta), "0", voz)
|
||||||
|
return len([x for x in salida.split() if x.strip().isdigit()])
|
||||||
|
|
||||||
|
|
||||||
|
def hechos_24h() -> dict[int, list[tuple[int, str]]]:
|
||||||
|
"""mp3 escritos en las últimas 24 h, agrupados por autor.
|
||||||
|
|
||||||
|
El mtime del fichero es la fuente: es lo que se acaba de escribir, sin
|
||||||
|
depender de metas que puedan venir de una sincronización antigua.
|
||||||
|
"""
|
||||||
|
corte = time.time() - 24 * 3600
|
||||||
|
por_autor: dict[int, list[tuple[int, str]]] = {}
|
||||||
|
if not TTS_DIR.is_dir():
|
||||||
|
return por_autor
|
||||||
|
recientes = [f for f in TTS_DIR.glob("*.mp3")
|
||||||
|
if f.stat().st_mtime >= corte and f.stem.isdigit()]
|
||||||
|
for f in sorted(recientes, key=lambda p: p.stat().st_mtime):
|
||||||
|
pid = int(f.stem)
|
||||||
|
# No hay meta de autor; la voz sí se guarda (fea_audio_voice) y basta
|
||||||
|
# para atribuirlo, porque cada autor clonado tiene la suya.
|
||||||
|
voz = php("getmeta", str(pid), "fea_audio_voice").strip()
|
||||||
|
aid = next((a for a, (_, v) in AUTORES.items() if v == voz), 0)
|
||||||
|
por_autor.setdefault(aid, []).append((pid, voz or "?"))
|
||||||
|
return por_autor
|
||||||
|
|
||||||
|
|
||||||
|
def cuota() -> tuple[int | None, int | None]:
|
||||||
|
try:
|
||||||
|
r = subprocess.run([sys.executable, str(QUOTA), "--json", "--no-local"],
|
||||||
|
capture_output=True, text=True, timeout=40)
|
||||||
|
d = json.loads(r.stdout)
|
||||||
|
m = next(p for p in d["providers"] if p["provider"] == "minimax" and p.get("ok"))
|
||||||
|
return m.get("five_h_pct"), m.get("week_pct")
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
return None, None
|
||||||
|
|
||||||
|
|
||||||
|
def logs_24h() -> tuple[int, int, list[str]]:
|
||||||
|
"""(ventanas ejecutadas, ventanas saltadas por gate, líneas de fallo)."""
|
||||||
|
hoy = datetime.now()
|
||||||
|
ficheros = [LOG_DIR / f"cron-{(hoy - timedelta(days=d)).strftime('%Y-%m-%d')}.log"
|
||||||
|
for d in (0, 1)]
|
||||||
|
corridas = saltadas = 0
|
||||||
|
fallos: list[str] = []
|
||||||
|
corte = hoy - timedelta(hours=24)
|
||||||
|
for f in ficheros:
|
||||||
|
if not f.is_file():
|
||||||
|
continue
|
||||||
|
for linea in f.read_text(errors="replace").splitlines():
|
||||||
|
m = re.match(r"\[(\d{4}-\d\d-\d\d \d\d:\d\d:\d\d)\]", linea)
|
||||||
|
if not m:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
if datetime.strptime(m.group(1), "%Y-%m-%d %H:%M:%S") < corte:
|
||||||
|
continue
|
||||||
|
except ValueError:
|
||||||
|
continue
|
||||||
|
if "cron TTS backlog start" in linea:
|
||||||
|
corridas += 1
|
||||||
|
elif "ABORT:" in linea:
|
||||||
|
saltadas += 1
|
||||||
|
fallos.append(linea.split("ABORT:", 1)[1].strip())
|
||||||
|
elif "FALLO rc=" in linea or "listpending falló" in linea:
|
||||||
|
fallos.append(linea.split("] ", 1)[-1].strip())
|
||||||
|
return corridas, saltadas, fallos
|
||||||
|
|
||||||
|
|
||||||
|
def _campo_cron(campo: str, valores: range) -> set[int]:
|
||||||
|
"""Expande un campo de crontab ('*', '*/5', '1,5,6,0', '0-4') a un set."""
|
||||||
|
if campo == "*":
|
||||||
|
return set(valores)
|
||||||
|
out: set[int] = set()
|
||||||
|
for trozo in campo.split(","):
|
||||||
|
paso = 1
|
||||||
|
if "/" in trozo:
|
||||||
|
trozo, p = trozo.split("/", 1)
|
||||||
|
paso = int(p)
|
||||||
|
if trozo == "*":
|
||||||
|
base = list(valores)
|
||||||
|
elif "-" in trozo:
|
||||||
|
a, b = (int(x) for x in trozo.split("-", 1))
|
||||||
|
base = list(range(a, b + 1))
|
||||||
|
else:
|
||||||
|
base = [int(trozo)]
|
||||||
|
out.update(base[::paso] if paso > 1 else base)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def previstas_24h(ahora: datetime | None = None) -> int | None:
|
||||||
|
"""Cuántas ventanas tenían que haber corrido en las últimas 24 horas.
|
||||||
|
|
||||||
|
El backlog usa varias líneas de crontab (ritmo normal + sábado + domingo).
|
||||||
|
Se calcula la unión de sus instantes previstos, sin contar dos veces una
|
||||||
|
coincidencia entre líneas.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
lineas = [
|
||||||
|
l for l in subprocess.run(["crontab", "-l"], text=True, capture_output=True,
|
||||||
|
check=True).stdout.splitlines()
|
||||||
|
if "tts_backlog_cron.sh" in l and not l.lstrip().startswith("#")
|
||||||
|
]
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
return None
|
||||||
|
if not lineas:
|
||||||
|
return None
|
||||||
|
|
||||||
|
ahora = ahora or datetime.now()
|
||||||
|
inicio = ahora - timedelta(hours=24)
|
||||||
|
previstos: set[datetime] = set()
|
||||||
|
for linea in lineas:
|
||||||
|
campos = linea.split(None, 5)
|
||||||
|
if len(campos) < 5:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
minutos = _campo_cron(campos[0], range(60))
|
||||||
|
horas = _campo_cron(campos[1], range(24))
|
||||||
|
dows = {d % 7 for d in _campo_cron(campos[4], range(7))}
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
continue
|
||||||
|
for h in range(25):
|
||||||
|
t = (ahora - timedelta(hours=h)).replace(second=0, microsecond=0)
|
||||||
|
for m in minutos:
|
||||||
|
cand = t.replace(minute=m)
|
||||||
|
if inicio < cand <= ahora and cand.hour in horas and (cand.weekday() + 1) % 7 in dows:
|
||||||
|
previstos.add(cand)
|
||||||
|
return len(previstos)
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
# Nunca informar «0 pendientes» cuando la lectura de WordPress ha fallado.
|
||||||
|
# El cron no-agent entrega stdout/errores tal cual: un aviso explícito permite
|
||||||
|
# arreglar Docker/helper sin que parezca que el backlog está terminado.
|
||||||
|
try:
|
||||||
|
hechos = hechos_24h()
|
||||||
|
faltantes = {
|
||||||
|
aid: (pendientes(aid, 0, 9999, voz),
|
||||||
|
pendientes(aid, ACTIVO[1], ACTIVO[2], voz) if aid == ACTIVO[0] else None)
|
||||||
|
for aid, (_nombre, voz) in AUTORES.items()
|
||||||
|
}
|
||||||
|
except RuntimeError as exc:
|
||||||
|
print("⚠️ Fe Adulta — informe TTS inválido: no se pudo leer WordPress.")
|
||||||
|
print(f" {exc}")
|
||||||
|
print(" Los ceros no son datos reales; revisar wordpress-web y /tmp/fea_post_io.php.")
|
||||||
|
return 1
|
||||||
|
|
||||||
|
total = sum(len(v) for v in hechos.values())
|
||||||
|
corridas, saltadas, fallos = logs_24h()
|
||||||
|
p5, pw = cuota()
|
||||||
|
|
||||||
|
lineas = [f"Fe Adulta — backlog TTS (últimas 24 h): {total} audios"]
|
||||||
|
|
||||||
|
if hechos:
|
||||||
|
for aid, items in sorted(hechos.items(), key=lambda kv: -len(kv[1])):
|
||||||
|
nombre = AUTORES.get(aid, ("otros", ""))[0]
|
||||||
|
lineas.append(f" {nombre}: {len(items)} "
|
||||||
|
f"({', '.join('#%d' % p for p, _ in items[:8])}"
|
||||||
|
f"{'…' if len(items) > 8 else ''})")
|
||||||
|
|
||||||
|
lineas.append("")
|
||||||
|
lineas.append("Pendientes:")
|
||||||
|
for aid, (nombre, voz) in AUTORES.items():
|
||||||
|
falta_todo, falta_lote = faltantes[aid]
|
||||||
|
marca = ""
|
||||||
|
if aid == ACTIVO[0]:
|
||||||
|
marca = f" ← en curso, {falta_lote} del lote {ACTIVO[1]}-{ACTIVO[2]}"
|
||||||
|
lineas.append(f" {nombre}: {falta_todo}{marca}")
|
||||||
|
|
||||||
|
lineas.append("")
|
||||||
|
if p5 is None:
|
||||||
|
lineas.append("Cuota MiniMax: no se pudo leer")
|
||||||
|
else:
|
||||||
|
lineas.append(f"Cuota MiniMax: 5h {p5:.0f}% · semana {pw:.0f}%")
|
||||||
|
previstas = previstas_24h()
|
||||||
|
de = "" if previstas is None else f" de {previstas} previstas"
|
||||||
|
lineas.append(f"Ventanas 24 h: {corridas} ejecutadas{de}, {saltadas} saltadas por cuota")
|
||||||
|
if previstas == 0:
|
||||||
|
lineas.append(" (sin ventanas previstas: día sin cron, toca carta)")
|
||||||
|
|
||||||
|
# Una ventana que ni arranca ni se salta es un fallo mudo: el cron no llegó a
|
||||||
|
# correr (bit +x, WSL apagada...). Es exactamente lo que pasó el 2-ago.
|
||||||
|
# Solo es alarma si de verdad tocaba correr; si no, es martes.
|
||||||
|
if corridas == 0 and saltadas == 0 and previstas != 0:
|
||||||
|
lineas.append("")
|
||||||
|
lineas.append("⚠️ Ninguna ventana dejó rastro en 24 h. Si tocaba que corriera, "
|
||||||
|
f"comprobar: crontab -l | grep tts_backlog · ls -l {CRON}")
|
||||||
|
|
||||||
|
if fallos:
|
||||||
|
lineas.append("")
|
||||||
|
lineas.append("Avisos:")
|
||||||
|
for f in fallos[:10]:
|
||||||
|
lineas.append(f" {f}")
|
||||||
|
|
||||||
|
print("\n".join(lineas))
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
Executable
+385
@@ -0,0 +1,385 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Query GA4 via the Google Analytics Data API and Admin API.
|
||||||
|
|
||||||
|
This script is intended for practical editorial analysis:
|
||||||
|
- resolve a GA4 property from a measurement ID
|
||||||
|
- run a few reusable reports
|
||||||
|
- export the result to CSV
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from google.auth.transport.requests import Request
|
||||||
|
from google.oauth2.credentials import Credentials
|
||||||
|
from google_auth_oauthlib.flow import InstalledAppFlow
|
||||||
|
|
||||||
|
SCOPES = ["https://www.googleapis.com/auth/analytics.readonly"]
|
||||||
|
DATA_API_BASE = "https://analyticsdata.googleapis.com/v1beta"
|
||||||
|
ADMIN_API_BASE = "https://analyticsadmin.googleapis.com/v1beta"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Config:
|
||||||
|
client_secrets_path: Path
|
||||||
|
token_path: Path
|
||||||
|
property_id: str | None
|
||||||
|
measurement_id: str | None
|
||||||
|
no_browser: bool
|
||||||
|
|
||||||
|
|
||||||
|
PRESETS: dict[str, dict[str, Any]] = {
|
||||||
|
"summary": {
|
||||||
|
"dimensions": [],
|
||||||
|
"metrics": ["screenPageViews", "totalUsers", "sessions", "engagedSessions", "engagementRate"],
|
||||||
|
"order_bys": [],
|
||||||
|
},
|
||||||
|
"traffic": {
|
||||||
|
"dimensions": ["date"],
|
||||||
|
"metrics": ["sessions", "totalUsers", "engagedSessions", "engagementRate", "screenPageViews"],
|
||||||
|
"order_bys": [{"dimension": {"dimensionName": "date"}}],
|
||||||
|
},
|
||||||
|
"content": {
|
||||||
|
"dimensions": ["pageTitle", "pagePath"],
|
||||||
|
"metrics": ["screenPageViews", "totalUsers", "engagedSessions", "engagementRate", "averageSessionDuration"],
|
||||||
|
"order_bys": [{"metric": {"metricName": "screenPageViews"}, "desc": True}],
|
||||||
|
},
|
||||||
|
"landing-pages": {
|
||||||
|
"dimensions": ["landingPagePlusQueryString"],
|
||||||
|
"metrics": ["sessions", "totalUsers", "engagedSessions", "engagementRate", "screenPageViews"],
|
||||||
|
"order_bys": [{"metric": {"metricName": "sessions"}, "desc": True}],
|
||||||
|
},
|
||||||
|
"source-medium": {
|
||||||
|
"dimensions": ["sessionSourceMedium"],
|
||||||
|
"metrics": ["sessions", "totalUsers", "engagedSessions", "engagementRate", "screenPageViews"],
|
||||||
|
"order_bys": [{"metric": {"metricName": "sessions"}, "desc": True}],
|
||||||
|
},
|
||||||
|
"device": {
|
||||||
|
"dimensions": ["deviceCategory"],
|
||||||
|
"metrics": ["sessions", "totalUsers", "engagedSessions", "engagementRate", "screenPageViews"],
|
||||||
|
"order_bys": [{"metric": {"metricName": "sessions"}, "desc": True}],
|
||||||
|
},
|
||||||
|
"hosts": {
|
||||||
|
"dimensions": ["hostName"],
|
||||||
|
"metrics": ["sessions", "totalUsers", "engagedSessions", "engagementRate", "screenPageViews"],
|
||||||
|
"order_bys": [{"metric": {"metricName": "sessions"}, "desc": True}],
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def load_config(args: argparse.Namespace) -> Config:
|
||||||
|
client_secrets = args.client_secrets_path or os.getenv("GA4_CLIENT_SECRETS_PATH")
|
||||||
|
token_path = args.token_path or os.getenv("GA4_TOKEN_PATH") or ".secrets/ga4-token.json"
|
||||||
|
property_id = args.property_id or os.getenv("GA4_PROPERTY_ID")
|
||||||
|
measurement_id = args.measurement_id or os.getenv("GA4_MEASUREMENT_ID")
|
||||||
|
|
||||||
|
if not client_secrets:
|
||||||
|
raise SystemExit(
|
||||||
|
"Missing OAuth client secrets path. Set --client-secrets-path or GA4_CLIENT_SECRETS_PATH."
|
||||||
|
)
|
||||||
|
|
||||||
|
return Config(
|
||||||
|
client_secrets_path=Path(client_secrets),
|
||||||
|
token_path=Path(token_path),
|
||||||
|
property_id=property_id,
|
||||||
|
measurement_id=measurement_id,
|
||||||
|
no_browser=bool(args.no_browser),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def get_credentials(config: Config) -> Credentials:
|
||||||
|
creds: Credentials | None = None
|
||||||
|
|
||||||
|
if config.token_path.exists():
|
||||||
|
creds = Credentials.from_authorized_user_file(str(config.token_path), SCOPES)
|
||||||
|
|
||||||
|
if creds and creds.valid:
|
||||||
|
return creds
|
||||||
|
|
||||||
|
if creds and creds.expired and creds.refresh_token:
|
||||||
|
creds.refresh(Request())
|
||||||
|
config.token_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
config.token_path.write_text(creds.to_json(), encoding="utf-8")
|
||||||
|
return creds
|
||||||
|
|
||||||
|
if not config.client_secrets_path.exists():
|
||||||
|
raise SystemExit(f"Client secrets file not found: {config.client_secrets_path}")
|
||||||
|
|
||||||
|
flow = InstalledAppFlow.from_client_secrets_file(str(config.client_secrets_path), SCOPES)
|
||||||
|
prompt_message = "Please visit this URL to authorize this application: {url}"
|
||||||
|
creds = flow.run_local_server(
|
||||||
|
port=0,
|
||||||
|
open_browser=not config.no_browser,
|
||||||
|
authorization_prompt_message=prompt_message,
|
||||||
|
)
|
||||||
|
config.token_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
config.token_path.write_text(creds.to_json(), encoding="utf-8")
|
||||||
|
return creds
|
||||||
|
|
||||||
|
|
||||||
|
def auth_headers(creds: Credentials) -> dict[str, str]:
|
||||||
|
if not creds.valid:
|
||||||
|
creds.refresh(Request())
|
||||||
|
return {
|
||||||
|
"Authorization": f"Bearer {creds.token}",
|
||||||
|
"Content-Type": "application/json",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def admin_get(creds: Credentials, path: str, params: dict[str, Any] | None = None) -> dict[str, Any]:
|
||||||
|
url = f"{ADMIN_API_BASE}/{path.lstrip('/')}"
|
||||||
|
response = requests.get(url, headers=auth_headers(creds), params=params, timeout=60)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
|
||||||
|
def data_post(creds: Credentials, path: str, payload: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
url = f"{DATA_API_BASE}/{path.lstrip('/')}"
|
||||||
|
response = requests.post(url, headers=auth_headers(creds), json=payload, timeout=60)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
|
||||||
|
def iterate_account_summaries(creds: Credentials) -> list[dict[str, Any]]:
|
||||||
|
results: list[dict[str, Any]] = []
|
||||||
|
page_token: str | None = None
|
||||||
|
|
||||||
|
while True:
|
||||||
|
params = {"pageSize": 200}
|
||||||
|
if page_token:
|
||||||
|
params["pageToken"] = page_token
|
||||||
|
payload = admin_get(creds, "accountSummaries", params=params)
|
||||||
|
results.extend(payload.get("accountSummaries", []))
|
||||||
|
page_token = payload.get("nextPageToken")
|
||||||
|
if not page_token:
|
||||||
|
return results
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_property_id(creds: Credentials, measurement_id: str) -> dict[str, str]:
|
||||||
|
summaries = iterate_account_summaries(creds)
|
||||||
|
|
||||||
|
for summary in summaries:
|
||||||
|
for prop in summary.get("propertySummaries", []):
|
||||||
|
prop_resource = prop.get("property", "")
|
||||||
|
if not prop_resource.startswith("properties/"):
|
||||||
|
continue
|
||||||
|
prop_id = prop_resource.split("/", 1)[1]
|
||||||
|
streams = admin_get(creds, f"properties/{prop_id}/dataStreams")
|
||||||
|
for stream in streams.get("dataStreams", []):
|
||||||
|
web_stream = stream.get("webStreamData", {})
|
||||||
|
if web_stream.get("measurementId") == measurement_id:
|
||||||
|
return {
|
||||||
|
"property_id": prop_id,
|
||||||
|
"property_display_name": prop.get("displayName", ""),
|
||||||
|
"account_display_name": summary.get("displayName", ""),
|
||||||
|
"stream_display_name": stream.get("displayName", ""),
|
||||||
|
}
|
||||||
|
|
||||||
|
raise SystemExit(f"No accessible GA4 property matched measurement ID {measurement_id}.")
|
||||||
|
|
||||||
|
|
||||||
|
def build_report_payload(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
preset = PRESETS[args.preset]
|
||||||
|
start_date = args.start_date or f"{args.days}daysAgo"
|
||||||
|
end_date = args.end_date or "yesterday"
|
||||||
|
payload: dict[str, Any] = {
|
||||||
|
"metrics": [{"name": m} for m in preset["metrics"]],
|
||||||
|
"dateRanges": [{"startDate": start_date, "endDate": end_date}],
|
||||||
|
"limit": str(args.limit),
|
||||||
|
"keepEmptyRows": False,
|
||||||
|
"returnPropertyQuota": True,
|
||||||
|
}
|
||||||
|
if preset["dimensions"]:
|
||||||
|
payload["dimensions"] = [{"name": d} for d in preset["dimensions"]]
|
||||||
|
if preset["order_bys"]:
|
||||||
|
payload["orderBys"] = preset["order_bys"]
|
||||||
|
filters: list[dict[str, Any]] = []
|
||||||
|
|
||||||
|
if args.page_path_regex:
|
||||||
|
expression: dict[str, Any] = {
|
||||||
|
"filter": {
|
||||||
|
"fieldName": "pagePath",
|
||||||
|
"stringFilter": {
|
||||||
|
"matchType": "FULL_REGEXP",
|
||||||
|
"value": args.page_path_regex,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if args.page_path_regex_not:
|
||||||
|
expression = {"notExpression": expression}
|
||||||
|
filters.append(expression)
|
||||||
|
|
||||||
|
# La propiedad G-6RT9ZRS4LW mide varios hostnames a la vez (www.feadulta.com
|
||||||
|
# vivo y antiguo.feadulta.com, el archivo estatico). Sin este filtro los
|
||||||
|
# informes los mezclan y no significan nada.
|
||||||
|
host_filter = getattr(args, "host", None)
|
||||||
|
if host_filter:
|
||||||
|
hosts = [h.strip() for h in host_filter.split(",") if h.strip()]
|
||||||
|
host_expression: dict[str, Any] = {
|
||||||
|
"filter": {
|
||||||
|
"fieldName": "hostName",
|
||||||
|
"inListFilter": {"values": hosts, "caseSensitive": False},
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if getattr(args, "host_not", False):
|
||||||
|
host_expression = {"notExpression": host_expression}
|
||||||
|
filters.append(host_expression)
|
||||||
|
|
||||||
|
if len(filters) == 1:
|
||||||
|
payload["dimensionFilter"] = filters[0]
|
||||||
|
elif len(filters) > 1:
|
||||||
|
payload["dimensionFilter"] = {"andGroup": {"expressions": filters}}
|
||||||
|
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def rows_from_response(response: dict[str, Any]) -> tuple[list[str], list[list[str]]]:
|
||||||
|
dimensions = [h["name"] for h in response.get("dimensionHeaders", [])]
|
||||||
|
metrics = [h["name"] for h in response.get("metricHeaders", [])]
|
||||||
|
headers = dimensions + metrics
|
||||||
|
rows: list[list[str]] = []
|
||||||
|
|
||||||
|
for row in response.get("rows", []):
|
||||||
|
dimension_values = [v.get("value", "") for v in row.get("dimensionValues", [])]
|
||||||
|
metric_values = [v.get("value", "") for v in row.get("metricValues", [])]
|
||||||
|
rows.append(dimension_values + metric_values)
|
||||||
|
|
||||||
|
return headers, rows
|
||||||
|
|
||||||
|
|
||||||
|
def write_csv(path: str, headers: list[str], rows: list[list[str]]) -> None:
|
||||||
|
out_path = Path(path)
|
||||||
|
out_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
with out_path.open("w", newline="", encoding="utf-8") as handle:
|
||||||
|
writer = csv.writer(handle)
|
||||||
|
writer.writerow(headers)
|
||||||
|
writer.writerows(rows)
|
||||||
|
|
||||||
|
|
||||||
|
def print_table(headers: list[str], rows: list[list[str]]) -> None:
|
||||||
|
widths = [len(h) for h in headers]
|
||||||
|
for row in rows:
|
||||||
|
for idx, value in enumerate(row):
|
||||||
|
widths[idx] = max(widths[idx], len(value))
|
||||||
|
|
||||||
|
fmt = " | ".join(f"{{:{w}}}" for w in widths)
|
||||||
|
print(fmt.format(*headers))
|
||||||
|
print("-+-".join("-" * w for w in widths))
|
||||||
|
for row in rows:
|
||||||
|
print(fmt.format(*row))
|
||||||
|
|
||||||
|
|
||||||
|
def cmd_resolve_property(args: argparse.Namespace) -> int:
|
||||||
|
config = load_config(args)
|
||||||
|
if not config.measurement_id:
|
||||||
|
raise SystemExit("Missing measurement ID. Set --measurement-id or GA4_MEASUREMENT_ID.")
|
||||||
|
|
||||||
|
creds = get_credentials(config)
|
||||||
|
result = resolve_property_id(creds, config.measurement_id)
|
||||||
|
print(json.dumps(result, indent=2, ensure_ascii=True))
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
def cmd_report(args: argparse.Namespace) -> int:
|
||||||
|
config = load_config(args)
|
||||||
|
creds = get_credentials(config)
|
||||||
|
|
||||||
|
property_id = config.property_id
|
||||||
|
if not property_id:
|
||||||
|
if not config.measurement_id:
|
||||||
|
raise SystemExit(
|
||||||
|
"Missing property ID. Set --property-id / GA4_PROPERTY_ID or provide --measurement-id / GA4_MEASUREMENT_ID."
|
||||||
|
)
|
||||||
|
resolved = resolve_property_id(creds, config.measurement_id)
|
||||||
|
property_id = resolved["property_id"]
|
||||||
|
print(
|
||||||
|
f"Resolved measurement ID {config.measurement_id} to property {property_id} "
|
||||||
|
f"({resolved['property_display_name']})",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
|
||||||
|
payload = build_report_payload(args)
|
||||||
|
response = data_post(creds, f"properties/{property_id}:runReport", payload)
|
||||||
|
headers, rows = rows_from_response(response)
|
||||||
|
|
||||||
|
if args.csv:
|
||||||
|
write_csv(args.csv, headers, rows)
|
||||||
|
print(f"Wrote CSV to {args.csv}", file=sys.stderr)
|
||||||
|
|
||||||
|
print_table(headers, rows)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = argparse.ArgumentParser(description="Query GA4 via OAuth.")
|
||||||
|
parser.add_argument("--client-secrets-path", help="Path to OAuth desktop client secrets JSON.")
|
||||||
|
parser.add_argument("--token-path", help="Path to cached OAuth token JSON.")
|
||||||
|
parser.add_argument("--property-id", help="GA4 property ID.")
|
||||||
|
parser.add_argument("--measurement-id", help="GA4 measurement ID (G-...).")
|
||||||
|
parser.add_argument(
|
||||||
|
"--no-browser",
|
||||||
|
action="store_true",
|
||||||
|
help="Print the OAuth URL instead of trying to open a browser automatically.",
|
||||||
|
)
|
||||||
|
|
||||||
|
subparsers = parser.add_subparsers(dest="command", required=True)
|
||||||
|
|
||||||
|
resolve_parser = subparsers.add_parser("resolve-property", help="Resolve GA4 property from measurement ID.")
|
||||||
|
resolve_parser.set_defaults(func=cmd_resolve_property)
|
||||||
|
|
||||||
|
report_parser = subparsers.add_parser("report", help="Run a preset GA4 report.")
|
||||||
|
report_parser.add_argument(
|
||||||
|
"--preset",
|
||||||
|
choices=sorted(PRESETS.keys()),
|
||||||
|
default="content",
|
||||||
|
help="Which report shape to run.",
|
||||||
|
)
|
||||||
|
report_parser.add_argument("--days", type=int, default=28, help="Lookback window in days.")
|
||||||
|
report_parser.add_argument("--start-date", help="Explicit GA4 start date, e.g. 2026-06-18.")
|
||||||
|
report_parser.add_argument("--end-date", help="Explicit GA4 end date, e.g. 2026-06-20.")
|
||||||
|
report_parser.add_argument("--limit", type=int, default=25, help="Max rows to request.")
|
||||||
|
report_parser.add_argument("--csv", help="Optional CSV output path.")
|
||||||
|
report_parser.add_argument(
|
||||||
|
"--page-path-regex",
|
||||||
|
help="Optional GA4 FULL_REGEXP filter applied to pagePath.",
|
||||||
|
)
|
||||||
|
report_parser.add_argument(
|
||||||
|
"--page-path-regex-not",
|
||||||
|
action="store_true",
|
||||||
|
help="Negate --page-path-regex.",
|
||||||
|
)
|
||||||
|
report_parser.add_argument(
|
||||||
|
"--host",
|
||||||
|
help=(
|
||||||
|
"Filtra por hostName (exacto, varios separados por coma). "
|
||||||
|
"Ej: www.feadulta.com o antiguo.feadulta.com. "
|
||||||
|
"Sin esto, la propiedad mezcla el sitio vivo y el archivo estatico."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
report_parser.add_argument(
|
||||||
|
"--host-not",
|
||||||
|
action="store_true",
|
||||||
|
help="Negate --host (todo MENOS esos hostnames).",
|
||||||
|
)
|
||||||
|
report_parser.set_defaults(func=cmd_report)
|
||||||
|
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
parser = build_parser()
|
||||||
|
args = parser.parse_args()
|
||||||
|
return args.func(args)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,135 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Importa las traducciones editoriales humanas adjuntas a un issue de carta.
|
||||||
|
|
||||||
|
Por defecto solo valida y muestra el plan. --apply-local crea borradores en el
|
||||||
|
WordPress Docker local; nunca publica ni toca producción.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import html
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
import tempfile
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
from xml.etree import ElementTree as ET
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parent
|
||||||
|
HELPER = ROOT / "fea_translate_helper.php"
|
||||||
|
CONTAINER = "wordpress-web"
|
||||||
|
SOURCE_ID = 55465 # Pagola, Carta 738
|
||||||
|
LANG_BY_NAME = {"2_eng": "en", "3_fr": "fr", "4_it": "it", "5_pt": "pt"}
|
||||||
|
TRANSLATOR_LABELS = ("Translator:", "Traducteur:", "Traduzzione:", "Tradutor:")
|
||||||
|
|
||||||
|
|
||||||
|
def docx_lines(path: Path) -> list[str]:
|
||||||
|
ns = {"w": "http://schemas.openxmlformats.org/wordprocessingml/2006/main"}
|
||||||
|
with zipfile.ZipFile(path) as zf:
|
||||||
|
root = ET.fromstring(zf.read("word/document.xml"))
|
||||||
|
return [
|
||||||
|
"".join(t.text or "" for t in p.findall(".//w:t", ns)).strip()
|
||||||
|
for p in root.findall(".//w:p", ns)
|
||||||
|
if "".join(t.text or "" for t in p.findall(".//w:t", ns)).strip()
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def value_after(lines: list[str], label: str) -> str:
|
||||||
|
try:
|
||||||
|
return lines[lines.index(label) + 1]
|
||||||
|
except (ValueError, IndexError) as exc:
|
||||||
|
raise ValueError(f"Falta {label!r}") from exc
|
||||||
|
|
||||||
|
|
||||||
|
def parse_doc(path: Path) -> dict:
|
||||||
|
lines = docx_lines(path)
|
||||||
|
title = value_after(lines, "Título:")
|
||||||
|
excerpt = value_after(lines, "Entradilla:")
|
||||||
|
author = value_after(lines, "Autor:")
|
||||||
|
start = lines.index("Cuerpo:") + 1
|
||||||
|
end = next((i for i in range(start, len(lines))
|
||||||
|
if any(lines[i].startswith(label) for label in TRANSLATOR_LABELS)), len(lines))
|
||||||
|
body = lines[start:end]
|
||||||
|
if not body:
|
||||||
|
raise ValueError("Cuerpo vacío")
|
||||||
|
translator = ""
|
||||||
|
if end < len(lines):
|
||||||
|
label = next(label for label in TRANSLATOR_LABELS if lines[end].startswith(label))
|
||||||
|
translator = lines[end][len(label):].strip()
|
||||||
|
if not translator and end + 1 < len(lines):
|
||||||
|
translator = lines[end + 1]
|
||||||
|
# La fuente española incluye el crédito/origen como último párrafo.
|
||||||
|
body.append('Publicado en: <a href="https://www.gruposdejesus.com">https://www.gruposdejesus.com</a>')
|
||||||
|
content = "\n".join(
|
||||||
|
f"<p>{line if line.startswith('Publicado en: <a ') else html.escape(line)}</p>" for line in body
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"title": title,
|
||||||
|
"excerpt": excerpt,
|
||||||
|
"author": author,
|
||||||
|
"translator": translator,
|
||||||
|
"content": content,
|
||||||
|
"source_sha256": hashlib.sha256(path.read_bytes()).hexdigest(),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def run(cmd: list[str], stdin: str | None = None) -> str:
|
||||||
|
result = subprocess.run(cmd, input=stdin, text=True, capture_output=True)
|
||||||
|
if result.returncode:
|
||||||
|
raise RuntimeError(f"rc={result.returncode}: {result.stderr.strip()} {result.stdout.strip()}")
|
||||||
|
return result.stdout.strip()
|
||||||
|
|
||||||
|
|
||||||
|
def helper(*args: str, stdin: str | None = None) -> str:
|
||||||
|
return run(["docker", "exec", "-i", CONTAINER, "php", "/tmp/fea_translate_helper.php", *args], stdin)
|
||||||
|
|
||||||
|
|
||||||
|
def lang_for(path: Path) -> str:
|
||||||
|
for marker, lang in LANG_BY_NAME.items():
|
||||||
|
if marker in path.name:
|
||||||
|
return lang
|
||||||
|
raise ValueError(f"No reconozco idioma en {path.name}")
|
||||||
|
|
||||||
|
|
||||||
|
def apply_meta(post_id: int, doc: Path, parsed: dict) -> None:
|
||||||
|
payload = json.dumps({"id": post_id, "doc": doc.name, "sha": parsed["source_sha256"],
|
||||||
|
"translator": parsed["translator"]}, ensure_ascii=False)
|
||||||
|
php = r'''$p=json_decode(base64_decode(getenv('PAYLOAD')),true); update_post_meta($p['id'],'traduccion_automatica','0'); update_post_meta($p['id'],'traduccion_modelo','editorial-humana'); update_post_meta($p['id'],'traduccion_fuente_doc',$p['doc']); update_post_meta($p['id'],'traduccion_fuente_sha256',$p['sha']); update_post_meta($p['id'],'traduccion_editor',$p['translator']); echo 'ok';'''
|
||||||
|
import base64
|
||||||
|
encoded = base64.b64encode(payload.encode()).decode()
|
||||||
|
run(["docker", "exec", "-i", "-e", f"PAYLOAD={encoded}", CONTAINER,
|
||||||
|
"wp", "--allow-root", "eval", php])
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument("--dir", type=Path, required=True)
|
||||||
|
ap.add_argument("--apply-local", action="store_true")
|
||||||
|
args = ap.parse_args()
|
||||||
|
docs = sorted(args.dir.glob("*.docx"))
|
||||||
|
if len(docs) != 4:
|
||||||
|
raise SystemExit(f"Se esperaban 4 DOCX, encontrados {len(docs)}")
|
||||||
|
run(["docker", "cp", str(HELPER), f"{CONTAINER}:/tmp/fea_translate_helper.php"])
|
||||||
|
for doc in docs:
|
||||||
|
lang, parsed = lang_for(doc), parse_doc(doc)
|
||||||
|
existing = helper("exists", str(SOURCE_ID), lang)
|
||||||
|
plan = {"lang": lang, "doc": doc.name, "existing": int(existing or "0"),
|
||||||
|
"title": parsed["title"], "excerpt_chars": len(parsed["excerpt"]),
|
||||||
|
"content_chars": len(parsed["content"]), "translator": parsed["translator"]}
|
||||||
|
print(json.dumps(plan, ensure_ascii=False))
|
||||||
|
if not args.apply_local:
|
||||||
|
continue
|
||||||
|
if int(existing or "0"):
|
||||||
|
raise RuntimeError(f"{lang} ya existe como #{existing}; no sobrescribo")
|
||||||
|
payload = json.dumps({"title": parsed["title"], "excerpt": parsed["excerpt"],
|
||||||
|
"content": parsed["content"], "model": "editorial-humana"}, ensure_ascii=False)
|
||||||
|
post_id = int(helper("create", str(SOURCE_ID), lang, "draft", stdin=payload))
|
||||||
|
apply_meta(post_id, doc, parsed)
|
||||||
|
print(json.dumps({"created": post_id, "lang": lang, "status": "draft"}))
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
+17
-7
@@ -93,7 +93,15 @@ def get_post_text(pid):
|
|||||||
check=True, capture_output=True)
|
check=True, capture_output=True)
|
||||||
subprocess.run(["docker", "cp", f"{CONTAINER}:/tmp/fea_es.json", "/tmp/fea_es.json"], check=True)
|
subprocess.run(["docker", "cp", f"{CONTAINER}:/tmp/fea_es.json", "/tmp/fea_es.json"], check=True)
|
||||||
d = json.load(open("/tmp/fea_es.json"))
|
d = json.load(open("/tmp/fea_es.json"))
|
||||||
raw = re.sub(r"(?i)</p>|<br\s*/?>|</h[1-6]>", "\n", d["content"])
|
# Hard gate: pages and operational/accounting entries are never TTS input.
|
||||||
|
# This also protects explicit --ids queues, which bypass author-backlog SQL.
|
||||||
|
raw_content = d.get("content", "")
|
||||||
|
blocked_markers = ("fea-don-wrap", "fea-ledger", "Haz tu donación")
|
||||||
|
if d.get("post_type") != "post":
|
||||||
|
raise ValueError(f"post #{pid} no es un artículo (post_type={d.get('post_type')!r}); TTS excluido")
|
||||||
|
if d.get("post_name") == "numeros" or any(marker in raw_content for marker in blocked_markers):
|
||||||
|
raise ValueError(f"post #{pid} es contenido de cuentas/donaciones; TTS excluido")
|
||||||
|
raw = re.sub(r"(?i)</p>|<br\s*/?>|</h[1-6]>", "\n", raw_content)
|
||||||
raw = re.sub(r"<[^>]+>", "", raw)
|
raw = re.sub(r"<[^>]+>", "", raw)
|
||||||
raw = re.sub(r"\[[^\]]+\]", "", raw)
|
raw = re.sub(r"\[[^\]]+\]", "", raw)
|
||||||
raw = html.unescape(raw)
|
raw = html.unescape(raw)
|
||||||
@@ -353,12 +361,12 @@ def _split_for_tts(text, limit=CHAR_LIMIT):
|
|||||||
return chunks
|
return chunks
|
||||||
|
|
||||||
|
|
||||||
def _synth_chunk(text, voice_id, model):
|
def _synth_chunk(text, voice_id, model, speed=1.0):
|
||||||
"""Una petición t2a. Devuelve (audio_bytes|None, rc, usage_chars)."""
|
"""Una petición t2a. Devuelve (audio_bytes|None, rc, usage_chars)."""
|
||||||
body = {
|
body = {
|
||||||
"model": model,
|
"model": model,
|
||||||
"text": text,
|
"text": text,
|
||||||
"voice_setting": {"voice_id": voice_id, "speed": 1.0, "vol": 1.0, "pitch": 0},
|
"voice_setting": {"voice_id": voice_id, "speed": speed, "vol": 1.0, "pitch": 0},
|
||||||
"audio_setting": {"sample_rate": 32000, "bitrate": 128000, "format": "mp3", "channel": 1},
|
"audio_setting": {"sample_rate": 32000, "bitrate": 128000, "format": "mp3", "channel": 1},
|
||||||
"language_boost": "Spanish",
|
"language_boost": "Spanish",
|
||||||
}
|
}
|
||||||
@@ -373,13 +381,15 @@ def _synth_chunk(text, voice_id, model):
|
|||||||
return bytes.fromhex(audio_hex), 0, usage
|
return bytes.fromhex(audio_hex), 0, usage
|
||||||
|
|
||||||
|
|
||||||
def t2a(text, voice_id, model, name):
|
def t2a(text, voice_id, model, name, speed=1.0):
|
||||||
|
if not 0.5 <= float(speed) <= 2.0:
|
||||||
|
raise ValueError(f"speed fuera de rango: {speed} (permitido 0.5–2.0)")
|
||||||
chunks = _split_for_tts(text)
|
chunks = _split_for_tts(text)
|
||||||
print(f"Sintetizando {len(text)} car con {model} / {voice_id} "
|
print(f"Sintetizando {len(text)} car con {model} / {voice_id} a {float(speed):.2f}× "
|
||||||
f"({len(chunks)} petición/es)…", flush=True)
|
f"({len(chunks)} petición/es)…", flush=True)
|
||||||
raw = OUT / f"{name}.raw.mp3"
|
raw = OUT / f"{name}.raw.mp3"
|
||||||
if len(chunks) == 1:
|
if len(chunks) == 1:
|
||||||
audio, rc, _ = _synth_chunk(chunks[0], voice_id, model)
|
audio, rc, _ = _synth_chunk(chunks[0], voice_id, model, speed)
|
||||||
if audio is None:
|
if audio is None:
|
||||||
return rc
|
return rc
|
||||||
raw.write_bytes(audio)
|
raw.write_bytes(audio)
|
||||||
@@ -390,7 +400,7 @@ def t2a(text, voice_id, model, name):
|
|||||||
import os as _os, time as _t
|
import os as _os, time as _t
|
||||||
_t.sleep(int(_os.environ.get("FEA_CHUNK_PAUSE", "35"))) # respetar TPM de MiniMax
|
_t.sleep(int(_os.environ.get("FEA_CHUNK_PAUSE", "35"))) # respetar TPM de MiniMax
|
||||||
print(f" trozo {k + 1}/{len(chunks)} ({len(ch)} car)…", flush=True)
|
print(f" trozo {k + 1}/{len(chunks)} ({len(ch)} car)…", flush=True)
|
||||||
audio, rc, _ = _synth_chunk(ch, voice_id, model)
|
audio, rc, _ = _synth_chunk(ch, voice_id, model, speed)
|
||||||
if audio is None:
|
if audio is None:
|
||||||
for p in parts:
|
for p in parts:
|
||||||
p.unlink(missing_ok=True)
|
p.unlink(missing_ok=True)
|
||||||
|
|||||||
@@ -0,0 +1,151 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Fase A local: borrador de Enrique + copia draft de la carta para revisión.
|
||||||
|
|
||||||
|
Por defecto no escribe WordPress. --apply-local sólo modifica el WordPress Docker local.
|
||||||
|
Nunca publica ni toca producción.
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import base64
|
||||||
|
import hashlib
|
||||||
|
import html
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
from xml.etree import ElementTree as ET
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parent.parent
|
||||||
|
DEFAULT_DOC = Path("/home/rafa/.hermes/cache/documents/doc_7909db597c74_7r_Enrique_Mart_nez_Lozano_tesoro_esta_ya_en_nosotros.docx")
|
||||||
|
CARTA_PROD_ID = 54914
|
||||||
|
AUTHOR_ID = 384 # Enrique Martínez Lozano, verificado localmente.
|
||||||
|
CATEGORIES = [1650, 71] # Artículos + Feadulta
|
||||||
|
SLUG = "el-tesoro-esta-ya-en-nosotros"
|
||||||
|
|
||||||
|
|
||||||
|
def docx_text(path: Path) -> list[str]:
|
||||||
|
ns = {"w": "http://schemas.openxmlformats.org/wordprocessingml/2006/main"}
|
||||||
|
with zipfile.ZipFile(path) as zf:
|
||||||
|
root = ET.fromstring(zf.read("word/document.xml"))
|
||||||
|
out = []
|
||||||
|
for p in root.findall(".//w:p", ns):
|
||||||
|
text = "".join(t.text or "" for t in p.findall(".//w:t", ns)).strip()
|
||||||
|
if text:
|
||||||
|
out.append(text)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def parse_source(path: Path) -> dict:
|
||||||
|
lines = docx_text(path)
|
||||||
|
def after(label: str) -> str:
|
||||||
|
return lines[lines.index(label) + 1]
|
||||||
|
title = after("Título:")
|
||||||
|
author = after("Autor:")
|
||||||
|
assert title == "EL TESORO ESTÁ YA EN NOSOTROS", title
|
||||||
|
assert author == "ENRIQUE MARTÍNEZ LOZANO", author
|
||||||
|
body_start = lines.index("Cuerpo:") + 1
|
||||||
|
body = lines[body_start:]
|
||||||
|
# Última firma del boletín no es contenido del post: el autor va en WP.
|
||||||
|
if body[-1].startswith("ENRIQUE MARTÍNEZ LOZANO"):
|
||||||
|
body = body[:-1]
|
||||||
|
date_line, bible, *paragraphs = body
|
||||||
|
intro = lines[lines.index("Entradilla:") + 1]
|
||||||
|
content = "\n".join(
|
||||||
|
[f"<p><strong>{html.escape(date_line)}</strong></p>",
|
||||||
|
f"<p><strong>{html.escape(bible)}</strong></p>"]
|
||||||
|
+ [f"<p>{html.escape(p)}</p>" for p in paragraphs]
|
||||||
|
)
|
||||||
|
return {"title": title.lower().capitalize(), "author": author.title(), "intro": intro,
|
||||||
|
"content": content, "source_sha256": hashlib.sha256(path.read_bytes()).hexdigest()}
|
||||||
|
|
||||||
|
|
||||||
|
def prod_card() -> dict:
|
||||||
|
sys.path.insert(0, str(ROOT / "scripts"))
|
||||||
|
import sync_translations_to_prod as sync
|
||||||
|
return json.loads(sync.prod_helper("read_full", str(CARTA_PROD_ID)))
|
||||||
|
|
||||||
|
|
||||||
|
def carta_candidate(card: dict, anchor: str) -> str:
|
||||||
|
marker = '<p> </p>\n<p><span style="color: #ff0000;"><strong>Para unas eucaristías más participativas y actuales</strong></span></p>'
|
||||||
|
assert card["content"].count(marker) == 1, "Marcador de sección no único/no encontrado"
|
||||||
|
return card["content"].replace(marker, anchor + "\n" + marker, 1)
|
||||||
|
|
||||||
|
|
||||||
|
def run_wp(payload: dict, apply: bool) -> dict:
|
||||||
|
encoded = base64.b64encode(json.dumps(payload, ensure_ascii=False).encode()).decode()
|
||||||
|
php = r'''
|
||||||
|
$p=json_decode(base64_decode(getenv('FEA_PAYLOAD')), true);
|
||||||
|
$apply=getenv('FEA_APPLY') === '1';
|
||||||
|
$existing=get_page_by_path($p['slug'], OBJECT, 'post');
|
||||||
|
$author=get_user_by('id',(int)$p['author_id']);
|
||||||
|
$cats=array_map('intval',$p['categories']);
|
||||||
|
$cat_ok=count(array_filter($cats, fn($id)=>term_exists($id,'category')))===count($cats);
|
||||||
|
$preview=get_posts(['post_type'=>'post','post_status'=>'any','meta_key'=>'fea_phase_a_source_hash','meta_value'=>$p['source_sha256'],'numberposts'=>1]);
|
||||||
|
$out=['dry_run'=>!$apply,'existing_slug'=>$existing?['id'=>$existing->ID,'status'=>$existing->post_status]:null,
|
||||||
|
'author'=>$author?['id'=>$author->ID,'name'=>$author->display_name]:null,'categories_ok'=>$cat_ok,
|
||||||
|
'preview_existing'=>$preview?['id'=>$preview[0]->ID,'status'=>$preview[0]->post_status]:null,
|
||||||
|
'candidate_anchor'=>$p['anchor']];
|
||||||
|
if (!$author || !$cat_ok) { $out['error']='Autor o categorías no válidos'; echo wp_json_encode($out,JSON_UNESCAPED_UNICODE); exit(3); }
|
||||||
|
if ($existing || $preview) { $out['error']='Ya existe un borrador/slug para esta Fase A; no se duplica'; echo wp_json_encode($out,JSON_UNESCAPED_UNICODE); exit(4); }
|
||||||
|
if (!$apply) { echo wp_json_encode($out,JSON_UNESCAPED_UNICODE); exit; }
|
||||||
|
$article_id=wp_insert_post(['post_type'=>'post','post_status'=>'draft','post_author'=>(int)$p['author_id'],
|
||||||
|
'post_title'=>$p['article_title'],'post_name'=>$p['slug'],'post_excerpt'=>$p['intro'],
|
||||||
|
'post_content'=>$p['article_content'],'post_category'=>$cats],true);
|
||||||
|
if (is_wp_error($article_id)) { $out['error']=$article_id->get_error_message(); echo wp_json_encode($out,JSON_UNESCAPED_UNICODE); exit(5); }
|
||||||
|
if (function_exists('pll_set_post_language')) pll_set_post_language($article_id,'es');
|
||||||
|
update_post_meta($article_id,'fea_phase_a_source_hash',$p['source_sha256']);
|
||||||
|
update_post_meta($article_id,'fea_phase_a_source_doc','7r_Enrique Martínez Lozano_tesoro_esta_ya_en_nosotros.docx');
|
||||||
|
update_post_meta($article_id,'fea_phase_a_status','preview-only');
|
||||||
|
$preview_id=wp_insert_post(['post_type'=>'post','post_status'=>'draft','post_author'=>(int)$p['card_author_id'],
|
||||||
|
'post_title'=>'PREVIEW — Hacia el corazón (+ Enrique Martínez Lozano)','post_excerpt'=>'Copia local de validación; no publicar.',
|
||||||
|
'post_content'=>$p['candidate_card_content']],true);
|
||||||
|
if (is_wp_error($preview_id)) { wp_delete_post($article_id,true); $out['error']=$preview_id->get_error_message(); echo wp_json_encode($out,JSON_UNESCAPED_UNICODE); exit(6); }
|
||||||
|
if (function_exists('pll_set_post_language')) pll_set_post_language($preview_id,'es');
|
||||||
|
update_post_meta($preview_id,'fea_phase_a_source_hash',$p['source_sha256']);
|
||||||
|
update_post_meta($preview_id,'fea_phase_a_source_prod_id',(int)$p['carta_prod_id']);
|
||||||
|
update_post_meta($preview_id,'fea_phase_a_preview_article_id',$article_id);
|
||||||
|
update_post_meta($preview_id,'fea_phase_a_status','preview-only');
|
||||||
|
$out['article']=['id'=>$article_id,'status'=>'draft','slug'=>get_post_field('post_name',$article_id),'author'=>get_the_author_meta('display_name',(int)$p['author_id'])];
|
||||||
|
$out['preview_card']=['id'=>$preview_id,'status'=>'draft','source_prod_id'=>(int)$p['carta_prod_id']];
|
||||||
|
echo wp_json_encode($out,JSON_UNESCAPED_UNICODE);
|
||||||
|
'''
|
||||||
|
cmd = [
|
||||||
|
"docker", "exec", "-i", "-e", f"FEA_PAYLOAD={encoded}",
|
||||||
|
"-e", f"FEA_APPLY={'1' if apply else '0'}", "wordpress-web",
|
||||||
|
"wp", "--allow-root", "eval", php,
|
||||||
|
]
|
||||||
|
result = subprocess.run(cmd, text=True, capture_output=True, check=False)
|
||||||
|
if result.returncode:
|
||||||
|
raise RuntimeError(f"WP local rc={result.returncode}: {result.stderr}\n{result.stdout}")
|
||||||
|
return json.loads(result.stdout)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument("--doc", type=Path, default=DEFAULT_DOC)
|
||||||
|
ap.add_argument("--apply-local", action="store_true")
|
||||||
|
ap.add_argument("--output", type=Path, required=True)
|
||||||
|
args = ap.parse_args()
|
||||||
|
source = parse_source(args.doc)
|
||||||
|
card = prod_card() # lectura server-side únicamente.
|
||||||
|
article_url = "http://localhost:8080/" + SLUG + "/"
|
||||||
|
anchor = (f'<p><strong><a href="{article_url}">Enrique Martínez Lozano: '
|
||||||
|
f'{html.escape(source["title"])}.</a></strong> {html.escape(source["intro"])}</p>')
|
||||||
|
payload = {
|
||||||
|
"source_sha256": source["source_sha256"], "article_title": source["title"],
|
||||||
|
"article_content": source["content"], "intro": source["intro"], "slug": SLUG,
|
||||||
|
"author_id": AUTHOR_ID, "categories": CATEGORIES, "anchor": anchor,
|
||||||
|
"candidate_card_content": carta_candidate(card, anchor), "card_author_id": card["author"],
|
||||||
|
"carta_prod_id": CARTA_PROD_ID,
|
||||||
|
}
|
||||||
|
result = run_wp(payload, args.apply_local)
|
||||||
|
report = {"source": source, "prod_card": {"id": card["id"], "title": card["title"], "status": card["status"]},
|
||||||
|
"result": result, "mode": "apply-local" if args.apply_local else "dry-run"}
|
||||||
|
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
args.output.write_text(json.dumps(report, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||||
|
print(json.dumps(report, ensure_ascii=False, indent=2))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
Executable
+167
@@ -0,0 +1,167 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Release Fase 1+2, issue #181: 68 traducciones publicadas + 16 MP3 TTS.
|
||||||
|
# Por defecto es sólo dry-run. --apply requiere confirmación explícita de Rafa.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||||
|
cd "$ROOT"
|
||||||
|
|
||||||
|
MODE="dry-run"
|
||||||
|
if [[ "${1:-}" == "--apply" ]]; then
|
||||||
|
MODE="apply"
|
||||||
|
elif [[ "${1:-}" == "--verify" ]]; then
|
||||||
|
MODE="verify"
|
||||||
|
elif [[ "${1:-}" != "" && "${1:-}" != "--dry-run" ]]; then
|
||||||
|
echo "Uso: $0 [--dry-run|--verify|--apply]" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
# shellcheck disable=SC1091
|
||||||
|
source ~/.hermes/profiles/feadulta/.env
|
||||||
|
: "${FEA_PROD_WPLOAD:=/web/wp-load.php}"
|
||||||
|
if [[ "$FEA_PROD_WPLOAD" != "/web/wp-load.php" ]]; then
|
||||||
|
echo "ABORT: FEA_PROD_WPLOAD debe ser /web/wp-load.php (recibido: $FEA_PROD_WPLOAD)" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
RELEASE_DIR="logs/release-issue-181"
|
||||||
|
mkdir -p "$RELEASE_DIR/backups"
|
||||||
|
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||||
|
ORIGINS=(54875 54902 54903 54883 54884 54885 54886 54866 54867 54868 54869 54870 54871 54872 54873 54880 54914)
|
||||||
|
AUDIO_IDS=(54875 54902 54903 54883 54884 54885 54886 54866 54867 54868 54869 54870 54871 54872 54873 54880)
|
||||||
|
|
||||||
|
if [[ "$MODE" == "apply" && "${FEA_RELEASE_CONFIRM:-}" != "PUBLISH_ISSUE_181" ]]; then
|
||||||
|
echo "ABORT: para escribir en producción usa:" >&2
|
||||||
|
echo " FEA_RELEASE_CONFIRM=PUBLISH_ISSUE_181 $0 --apply" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
preflight() {
|
||||||
|
python3 - "${ORIGINS[@]}" <<'PY'
|
||||||
|
import json, sys
|
||||||
|
sys.path.insert(0, 'scripts')
|
||||||
|
import sync_translations_to_prod as sync
|
||||||
|
ids = [int(x) for x in sys.argv[1:]]
|
||||||
|
rows = []
|
||||||
|
for pid in ids:
|
||||||
|
d = json.loads(sync.prod_helper('read_full', str(pid)))
|
||||||
|
rows.append({'id': pid, 'status': d['status'], 'lang': d['lang'], 'translations': d.get('translations', {})})
|
||||||
|
assert len(rows) == 17
|
||||||
|
assert all(r['status'] == 'publish' and r['lang'] == 'es' for r in rows), rows
|
||||||
|
# Los 16 artículos no deben tener traducciones aún; la carta tampoco antes del primer release.
|
||||||
|
assert all(r['translations'] == {'es': r['id']} for r in rows), rows
|
||||||
|
print(json.dumps(rows, ensure_ascii=False, indent=2))
|
||||||
|
print('preflight=PASS sources=17')
|
||||||
|
PY
|
||||||
|
}
|
||||||
|
|
||||||
|
backup_prod_state() {
|
||||||
|
python3 - "$RELEASE_DIR/backups/prod-before-$STAMP.json" "${ORIGINS[@]}" -- "${AUDIO_IDS[@]}" <<'PY'
|
||||||
|
import json, sys
|
||||||
|
out = sys.argv[1]
|
||||||
|
sep = sys.argv.index('--')
|
||||||
|
origins = [int(x) for x in sys.argv[2:sep]]
|
||||||
|
audio_ids = [int(x) for x in sys.argv[sep+1:]]
|
||||||
|
sys.path.insert(0, 'scripts')
|
||||||
|
import sync_translations_to_prod as translations
|
||||||
|
import sync_audio_to_prod as audio
|
||||||
|
payload = {
|
||||||
|
'origins': {str(pid): json.loads(translations.prod_helper('read_full', str(pid))) for pid in origins},
|
||||||
|
'audio_before': {
|
||||||
|
str(pid): {
|
||||||
|
'url': audio.prod_helper('getmeta', str(pid), 'fea_audio_url').strip(),
|
||||||
|
'voice': audio.prod_helper('getmeta', str(pid), 'fea_audio_voice').strip(),
|
||||||
|
'done': audio.prod_helper('getmeta', str(pid), 'fea_audio_done').strip(),
|
||||||
|
} for pid in audio_ids
|
||||||
|
},
|
||||||
|
}
|
||||||
|
with open(out, 'w', encoding='utf-8') as fh:
|
||||||
|
json.dump(payload, fh, ensure_ascii=False, indent=2)
|
||||||
|
print(out)
|
||||||
|
PY
|
||||||
|
}
|
||||||
|
|
||||||
|
verify_release() {
|
||||||
|
python3 - "${ORIGINS[@]}" -- "${AUDIO_IDS[@]}" <<'PY'
|
||||||
|
import json, sys, time
|
||||||
|
sep = sys.argv.index('--')
|
||||||
|
origins = [int(x) for x in sys.argv[1:sep]]
|
||||||
|
audio_ids = [int(x) for x in sys.argv[sep+1:]]
|
||||||
|
sys.path.insert(0, 'scripts')
|
||||||
|
import sync_translations_to_prod as translations
|
||||||
|
import sync_audio_to_prod as audio
|
||||||
|
|
||||||
|
def read_post(pid):
|
||||||
|
"""Lectura server-side resiliente ante una respuesta SSH/PHP vacía transitoria."""
|
||||||
|
last = ''
|
||||||
|
for attempt in range(1, 4):
|
||||||
|
raw = translations.prod_helper('read_full', str(pid)).strip()
|
||||||
|
if raw:
|
||||||
|
try:
|
||||||
|
return json.loads(raw)
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
last = f'JSON inválido intento {attempt}: {exc}; prefijo={raw[:160]!r}'
|
||||||
|
else:
|
||||||
|
last = f'respuesta vacía intento {attempt}'
|
||||||
|
time.sleep(attempt)
|
||||||
|
raise AssertionError(f'No se pudo leer post #{pid}: {last}')
|
||||||
|
|
||||||
|
for pid in origins:
|
||||||
|
d = read_post(pid)
|
||||||
|
group = d.get('translations', {})
|
||||||
|
assert set(group) == {'es', 'en', 'fr', 'it', 'pt'}, (pid, group)
|
||||||
|
for lang, tid in group.items():
|
||||||
|
td = read_post(tid)
|
||||||
|
assert td['lang'] == lang and td['status'] == 'publish', (pid, lang, tid, td['lang'], td['status'])
|
||||||
|
for pid in audio_ids:
|
||||||
|
url = audio.prod_helper('getmeta', str(pid), 'fea_audio_url').strip()
|
||||||
|
done = audio.prod_helper('getmeta', str(pid), 'fea_audio_done').strip()
|
||||||
|
assert url.endswith(f'/wp-content/uploads/tts/{pid}.mp3') and done == '1', (pid, url, done)
|
||||||
|
print('verification=PASS groups=17 translations=68 audio=16')
|
||||||
|
PY
|
||||||
|
}
|
||||||
|
|
||||||
|
if [[ "$MODE" == "verify" ]]; then
|
||||||
|
echo "== Verificación server-side posterior (solo lectura) =="
|
||||||
|
verify_release | tee "$RELEASE_DIR/verification-$STAMP.txt"
|
||||||
|
echo "VERIFICACIÓN ISSUE #181 COMPLETADA"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "== Preflight server-side ($MODE): 17 fuentes ES =="
|
||||||
|
preflight | tee "$RELEASE_DIR/preflight-$MODE-$STAMP.json"
|
||||||
|
|
||||||
|
if [[ "$MODE" == "dry-run" ]]; then
|
||||||
|
: > "$RELEASE_DIR/translations-dry-run-$STAMP.log"
|
||||||
|
for origin in "${ORIGINS[@]}"; do
|
||||||
|
FEA_SYNC_STATUS=publish \
|
||||||
|
FEA_SYNC_LOG="$RELEASE_DIR/translations-dry-run-$STAMP.log" \
|
||||||
|
FEA_SYNC_STATE="$RELEASE_DIR/translations-$origin-state.json" \
|
||||||
|
python3 scripts/sync_translations_to_prod.py --origin "$origin" --dry-run
|
||||||
|
done
|
||||||
|
FEA_AUDIO_SYNC_LOG="$RELEASE_DIR/audio-dry-run-$STAMP.log" \
|
||||||
|
FEA_AUDIO_SYNC_STATE="$RELEASE_DIR/audio-state.json" \
|
||||||
|
python3 scripts/sync_audio_to_prod.py --ids "$(IFS=,; echo "${AUDIO_IDS[*]}")" --dry-run | tee "$RELEASE_DIR/audio-dry-run-$STAMP.log"
|
||||||
|
echo "DRY-RUN terminado: 68 traducciones publicables + 16 audios planificados."
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "== Backup server-side previo =="
|
||||||
|
backup_prod_state | tee "$RELEASE_DIR/backup-path-$STAMP.txt"
|
||||||
|
|
||||||
|
echo "== Subiendo 68 traducciones como publish =="
|
||||||
|
for origin in "${ORIGINS[@]}"; do
|
||||||
|
FEA_SYNC_STATUS=publish \
|
||||||
|
FEA_SYNC_LOG="$RELEASE_DIR/translations-apply-$STAMP.log" \
|
||||||
|
FEA_SYNC_STATE="$RELEASE_DIR/translations-$origin-state.json" \
|
||||||
|
python3 scripts/sync_translations_to_prod.py --origin "$origin"
|
||||||
|
done
|
||||||
|
|
||||||
|
echo "== Subiendo 16 audios =="
|
||||||
|
FEA_AUDIO_SYNC_LOG="$RELEASE_DIR/audio-apply-$STAMP.log" \
|
||||||
|
FEA_AUDIO_SYNC_STATE="$RELEASE_DIR/audio-state.json" \
|
||||||
|
python3 scripts/sync_audio_to_prod.py --ids "$(IFS=,; echo "${AUDIO_IDS[*]}")"
|
||||||
|
|
||||||
|
echo "== Verificación server-side posterior =="
|
||||||
|
verify_release | tee "$RELEASE_DIR/verification-$STAMP.txt"
|
||||||
|
echo "RELEASE ISSUE #181 COMPLETADO"
|
||||||
+101
-39
@@ -3,9 +3,10 @@
|
|||||||
sync_audio_to_prod.py — Sube a PROD los mp3 de TTS ya generados/enlazados en
|
sync_audio_to_prod.py — Sube a PROD los mp3 de TTS ya generados/enlazados en
|
||||||
local (fea_audio_done=1) y fija el meta fea_audio_url en prod.
|
local (fea_audio_done=1) y fija el meta fea_audio_url en prod.
|
||||||
|
|
||||||
Prod (134.0.10.170) tiene glibc rota: scp/sftp NO funcionan (connection closed).
|
Prod vive en Hetzner (Coolify/Docker) desde el cutover de agosto 2026. Subida
|
||||||
Workaround: subir el binario por stdin de ssh ("cat > ruta"), igual que el resto
|
del binario por stdin de ssh ("cat > ruta" dentro del contenedor vía
|
||||||
de scripts que tocan ese servidor (ver feadulta-server-glibc-rota.md).
|
`docker exec -i`), sin depender de scp/sftp — mismo patrón que
|
||||||
|
sync_translations_to_prod.py y sync_carta_from_prod.py.
|
||||||
|
|
||||||
Uso:
|
Uso:
|
||||||
python3 sync_audio_to_prod.py --carta 54254 # sincroniza toda la cola de la carta
|
python3 sync_audio_to_prod.py --carta 54254 # sincroniza toda la cola de la carta
|
||||||
@@ -19,8 +20,10 @@ Rollback (despublica en prod lo que este script publicó):
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import argparse
|
import argparse
|
||||||
|
import base64
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
|
import shlex
|
||||||
import subprocess
|
import subprocess
|
||||||
import time
|
import time
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
@@ -31,13 +34,20 @@ DB_NAME = os.environ.get("FEA_DB_NAME", "wordpress_db")
|
|||||||
DB_USER = os.environ.get("FEA_DB_USER", "wordpress_user")
|
DB_USER = os.environ.get("FEA_DB_USER", "wordpress_user")
|
||||||
DB_PASS = os.environ.get("FEA_DB_PASS", "wordpress_pass")
|
DB_PASS = os.environ.get("FEA_DB_PASS", "wordpress_pass")
|
||||||
|
|
||||||
PROD_HOST = os.environ.get("FEA_PROD_HOST", "feadulta@134.0.10.170")
|
PROD_HOST = os.environ.get("FEA_PROD_SSH_HOST", "")
|
||||||
PROD_PASS = os.environ.get("FEA_PROD_PASS", "C6c2A!mAl3Wj.BQF")
|
PROD_PASS = os.environ.get("FEA_PROD_SSH_PASS", "")
|
||||||
PROD_WPLOAD = os.environ.get("FEA_PROD_WPLOAD", "/web/wp-load.php")
|
PROD_WPLOAD = os.environ.get("FEA_PROD_WPLOAD", "/var/www/html/wp-load.php")
|
||||||
PROD_HELPER = "/tmp/fea_post_io.php"
|
# Desde el cutover a Hetzner, WordPress vive dentro de Coolify/Docker. Si se
|
||||||
PROD_UPLOADS_TTS = "/web/wp-content/uploads/tts"
|
# define, el helper y la subida de mp3 se ejecutan dentro del contenedor.
|
||||||
|
# Espejo de sync_translations_to_prod.py / sync_carta_from_prod.py.
|
||||||
|
PROD_DOCKER_CONTAINER = os.environ.get("FEA_PROD_DOCKER_CONTAINER", "")
|
||||||
|
PROD_UPLOADS_TTS = os.environ.get("FEA_PROD_UPLOADS_TTS", "/var/www/html/wp-content/uploads/tts")
|
||||||
|
|
||||||
HELPER_SRC = Path(__file__).resolve().parent / "fea_post_io.php"
|
HELPER_SRC = Path(__file__).resolve().parent / "fea_post_io.php"
|
||||||
|
# Se evalúa en memoria con `php -r`: no se copia un helper temporal a prod.
|
||||||
|
HELPER_EVAL_CODE = "eval(base64_decode(" + repr(
|
||||||
|
base64.b64encode(HELPER_SRC.read_text(encoding="utf-8").removeprefix("<?php").encode("utf-8")).decode("ascii")
|
||||||
|
) + "));"
|
||||||
LOCAL_TTS_DIR = Path(__file__).resolve().parent.parent / "wordpress/wp-content/uploads/tts"
|
LOCAL_TTS_DIR = Path(__file__).resolve().parent.parent / "wordpress/wp-content/uploads/tts"
|
||||||
|
|
||||||
LOG_FILE = Path(os.environ.get(
|
LOG_FILE = Path(os.environ.get(
|
||||||
@@ -74,6 +84,17 @@ def local_meta(post_id: int, key: str) -> str:
|
|||||||
return r.stdout.strip()
|
return r.stdout.strip()
|
||||||
|
|
||||||
|
|
||||||
|
def prod_id_for_local(local_id: int) -> int:
|
||||||
|
"""Resolve a remapped local post to its original production ID.
|
||||||
|
|
||||||
|
``sync_carta_from_prod.py --remap-conflicts`` keeps this mapping in
|
||||||
|
``fea_prod_source_id``. Audio files remain named with the local ID, but
|
||||||
|
remote paths and production post meta must use the source ID.
|
||||||
|
"""
|
||||||
|
source_id = local_meta(local_id, "fea_prod_source_id")
|
||||||
|
return int(source_id) if source_id.isdigit() else local_id
|
||||||
|
|
||||||
|
|
||||||
def carta_article_ids(carta_id: int) -> list[int]:
|
def carta_article_ids(carta_id: int) -> list[int]:
|
||||||
q = ("SELECT post_id FROM wp_postmeta "
|
q = ("SELECT post_id FROM wp_postmeta "
|
||||||
f"WHERE meta_key='_carta_id' AND meta_value='{carta_id}' ORDER BY post_id;")
|
f"WHERE meta_key='_carta_id' AND meta_value='{carta_id}' ORDER BY post_id;")
|
||||||
@@ -85,47 +106,80 @@ def carta_article_ids(carta_id: int) -> list[int]:
|
|||||||
return [int(x) for x in r.stdout.split() if x.isdigit()]
|
return [int(x) for x in r.stdout.split() if x.isdigit()]
|
||||||
|
|
||||||
|
|
||||||
# ── Prod (glibc rota: nada de scp/sftp, todo por ssh + cat) ────────────────────
|
# ── Prod (Hetzner/Docker: todo por ssh, mp3 por stdin de "cat") ────────────────
|
||||||
def _ssh_text(remote_cmd: str, *, stdin: str | None = None, timeout: int = 120) -> str:
|
def _ssh_text(remote_cmd: str, *, stdin: str | None = None, timeout: int = 120) -> str:
|
||||||
|
# El Hetzner nuevo usa auth por clave (ed25519 claude-code@feadulta); sshpass
|
||||||
|
# solo se antepone si hay contraseña configurada (servidor viejo CDMON).
|
||||||
|
if PROD_PASS:
|
||||||
cmd = ["sshpass", "-p", PROD_PASS, "ssh", "-o", "StrictHostKeyChecking=accept-new",
|
cmd = ["sshpass", "-p", PROD_PASS, "ssh", "-o", "StrictHostKeyChecking=accept-new",
|
||||||
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
||||||
|
else:
|
||||||
|
cmd = ["ssh", "-o", "StrictHostKeyChecking=accept-new",
|
||||||
|
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
||||||
r = subprocess.run(cmd, input=stdin, capture_output=True, text=True, timeout=timeout)
|
r = subprocess.run(cmd, input=stdin, capture_output=True, text=True, timeout=timeout)
|
||||||
if r.returncode != 0:
|
if r.returncode != 0:
|
||||||
raise RuntimeError(f"ssh falló ({r.returncode}): {remote_cmd[:80]}…\n{r.stderr.strip()[:400]}")
|
raise RuntimeError(f"ssh falló ({r.returncode}): {remote_cmd[:80]}…\n{r.stderr.strip()[:400]}")
|
||||||
return r.stdout
|
return r.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def _remote_wrap(inner_cmd: str) -> str:
|
||||||
|
"""Envuelve un comando para que corra dentro del contenedor Docker de prod
|
||||||
|
si FEA_PROD_DOCKER_CONTAINER está definido; si no, corre en el host tal cual.
|
||||||
|
|
||||||
|
Importante: el wrapping (incluidas redirecciones como `< ruta`) debe quedar
|
||||||
|
DENTRO de la shell del contenedor (`sh -c '...'`), porque `docker exec
|
||||||
|
CONTENEDOR cmd < ruta` resuelve esa redirección en el filesystem del HOST,
|
||||||
|
no dentro del contenedor.
|
||||||
|
"""
|
||||||
|
if PROD_DOCKER_CONTAINER:
|
||||||
|
return f"docker exec -i {shlex.quote(PROD_DOCKER_CONTAINER)} sh -c {shlex.quote(inner_cmd)}"
|
||||||
|
return inner_cmd
|
||||||
|
|
||||||
|
|
||||||
def _ssh_upload_bytes(data: bytes, remote_path: str, *, timeout: int = 180) -> None:
|
def _ssh_upload_bytes(data: bytes, remote_path: str, *, timeout: int = 180) -> None:
|
||||||
|
remote_cmd = _remote_wrap(f"cat > {remote_path}")
|
||||||
|
if PROD_PASS:
|
||||||
cmd = ["sshpass", "-p", PROD_PASS, "ssh", "-o", "StrictHostKeyChecking=accept-new",
|
cmd = ["sshpass", "-p", PROD_PASS, "ssh", "-o", "StrictHostKeyChecking=accept-new",
|
||||||
"-o", "ConnectTimeout=20", PROD_HOST, f"cat > {remote_path}"]
|
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
||||||
|
else:
|
||||||
|
cmd = ["ssh", "-o", "StrictHostKeyChecking=accept-new",
|
||||||
|
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
||||||
sh(cmd, input_bytes=data, timeout=timeout)
|
sh(cmd, input_bytes=data, timeout=timeout)
|
||||||
|
|
||||||
|
|
||||||
_prod_helper_ready = False
|
|
||||||
|
|
||||||
|
|
||||||
def prod_helper(subcmd: str, *args: str) -> str:
|
def prod_helper(subcmd: str, *args: str) -> str:
|
||||||
global _prod_helper_ready
|
"""Ejecuta el helper PHP en memoria; no deja ficheros temporales en prod."""
|
||||||
if not _prod_helper_ready:
|
argv = " ".join(shlex.quote(x) for x in (subcmd, *args))
|
||||||
_ssh_upload_bytes(HELPER_SRC.read_bytes(), PROD_HELPER)
|
inner = (
|
||||||
_prod_helper_ready = True
|
f"FEA_WP_LOAD={shlex.quote(PROD_WPLOAD)} "
|
||||||
inner = f"FEA_WP_LOAD={PROD_WPLOAD} php {PROD_HELPER} {subcmd} " + " ".join(args)
|
f"php -r {shlex.quote(HELPER_EVAL_CODE)} {argv}"
|
||||||
return _ssh_text(inner, timeout=60)
|
)
|
||||||
|
if PROD_DOCKER_CONTAINER:
|
||||||
|
remote_cmd = (
|
||||||
|
f"docker exec -i -e FEA_WP_LOAD={shlex.quote(PROD_WPLOAD)} "
|
||||||
|
f"{shlex.quote(PROD_DOCKER_CONTAINER)} php -r {shlex.quote(HELPER_EVAL_CODE)} {argv}"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
remote_cmd = inner
|
||||||
|
return _ssh_text(remote_cmd, timeout=60)
|
||||||
|
|
||||||
|
|
||||||
def prod_upload_mp3(post_id: int) -> None:
|
def prod_upload_mp3(local_id: int, prod_id: int) -> None:
|
||||||
src = LOCAL_TTS_DIR / f"{post_id}.mp3"
|
src = LOCAL_TTS_DIR / f"{local_id}.mp3"
|
||||||
data = src.read_bytes()
|
data = src.read_bytes()
|
||||||
remote_path = f"{PROD_UPLOADS_TTS}/{post_id}.mp3"
|
remote_path = f"{PROD_UPLOADS_TTS}/{prod_id}.mp3"
|
||||||
_ssh_upload_bytes(data, remote_path)
|
_ssh_upload_bytes(data, remote_path)
|
||||||
# Verificación de tamaño (glibc rota => sin fiarse ciegamente del rc=0 de ssh)
|
# Verificación de tamaño: no fiarse ciegamente del rc=0 de ssh. `wc -c` debe
|
||||||
remote_size = int(_ssh_text(f"wc -c < {remote_path}").strip())
|
# correr (y resolver la redirección) DENTRO del contenedor — ver _remote_wrap.
|
||||||
|
remote_size = int(_ssh_text(_remote_wrap(f"wc -c < {remote_path}")).strip())
|
||||||
if remote_size != len(data):
|
if remote_size != len(data):
|
||||||
raise RuntimeError(f"tamaño no coincide tras subir #{post_id}: local={len(data)} remoto={remote_size}")
|
raise RuntimeError(
|
||||||
|
f"tamaño no coincide tras subir prod#{prod_id}: local={len(data)} remoto={remote_size}"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def prod_remove_mp3(post_id: int) -> None:
|
def prod_remove_mp3(post_id: int) -> None:
|
||||||
_ssh_text(f"rm -f {PROD_UPLOADS_TTS}/{post_id}.mp3")
|
_ssh_text(_remote_wrap(f"rm -f {PROD_UPLOADS_TTS}/{post_id}.mp3"))
|
||||||
|
|
||||||
|
|
||||||
# ── Estado ───────────────────────────────────────────────────────────────────
|
# ── Estado ───────────────────────────────────────────────────────────────────
|
||||||
@@ -149,25 +203,27 @@ def sync_one(post_id: int, state: dict, *, dry_run: bool) -> str:
|
|||||||
return "sin-audio-local"
|
return "sin-audio-local"
|
||||||
if not (LOCAL_TTS_DIR / f"{post_id}.mp3").exists():
|
if not (LOCAL_TTS_DIR / f"{post_id}.mp3").exists():
|
||||||
return "mp3-local-ausente"
|
return "mp3-local-ausente"
|
||||||
|
prod_id = prod_id_for_local(post_id)
|
||||||
if dry_run:
|
if dry_run:
|
||||||
return "PLAN: subiría mp3 + setaudio"
|
return f"PLAN: local#{post_id}.mp3 → prod#{prod_id}.mp3 + setaudio"
|
||||||
|
|
||||||
voice = local_meta(post_id, "fea_audio_voice") or "NicoFeadulta2026"
|
voice = local_meta(post_id, "fea_audio_voice") or "NicoFeadulta2026"
|
||||||
prod_upload_mp3(post_id)
|
prod_upload_mp3(post_id, prod_id)
|
||||||
prod_helper("setaudio", str(post_id), f"/wp-content/uploads/tts/{post_id}.mp3", voice)
|
prod_helper("setaudio", str(prod_id), f"/wp-content/uploads/tts/{prod_id}.mp3", voice)
|
||||||
if post_id not in state["synced"]:
|
if prod_id not in state["synced"]:
|
||||||
state["synced"].append(post_id)
|
state["synced"].append(prod_id)
|
||||||
save_state(state)
|
save_state(state)
|
||||||
return "ok"
|
return f"ok prod#{prod_id}"
|
||||||
|
|
||||||
|
|
||||||
def rollback_one(post_id: int, state: dict) -> str:
|
def rollback_one(post_id: int, state: dict) -> str:
|
||||||
prod_helper("unsetaudio", str(post_id))
|
prod_id = prod_id_for_local(post_id)
|
||||||
prod_remove_mp3(post_id)
|
prod_helper("unsetaudio", str(prod_id))
|
||||||
if post_id in state["synced"]:
|
prod_remove_mp3(prod_id)
|
||||||
state["synced"].remove(post_id)
|
if prod_id in state["synced"]:
|
||||||
|
state["synced"].remove(prod_id)
|
||||||
save_state(state)
|
save_state(state)
|
||||||
return "ok"
|
return f"ok prod#{prod_id}"
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
def main() -> int:
|
||||||
@@ -194,7 +250,7 @@ def main() -> int:
|
|||||||
for pid in ids:
|
for pid in ids:
|
||||||
try:
|
try:
|
||||||
if args.rollback:
|
if args.rollback:
|
||||||
if pid not in state["synced"] and not args.ids:
|
if prod_id_for_local(pid) not in state["synced"] and not args.ids:
|
||||||
res = "no-estaba-sincronizado"
|
res = "no-estaba-sincronizado"
|
||||||
skip += 1
|
skip += 1
|
||||||
else:
|
else:
|
||||||
@@ -202,7 +258,13 @@ def main() -> int:
|
|||||||
ok += 1
|
ok += 1
|
||||||
else:
|
else:
|
||||||
res = sync_one(pid, state, dry_run=args.dry_run)
|
res = sync_one(pid, state, dry_run=args.dry_run)
|
||||||
if res == "ok" or res.startswith("PLAN"):
|
# `startswith`, no `==`: sync_one devuelve "ok prod#<id>" desde que
|
||||||
|
# el id de prod puede diferir del local. Con la igualdad exacta, una
|
||||||
|
# subida perfecta se contaba entera como saltada — el 8-ago-2026 el
|
||||||
|
# pie del log dijo "ok=0 skip=90" tras subir los 90 audios de Fray
|
||||||
|
# Marcos sin un solo fallo. Peor que el susto: así un fallo real se
|
||||||
|
# confunde con este ruido y pasa desapercibido.
|
||||||
|
if res.startswith(("ok", "PLAN")):
|
||||||
ok += 1
|
ok += 1
|
||||||
else:
|
else:
|
||||||
skip += 1
|
skip += 1
|
||||||
|
|||||||
+100
-33
@@ -23,6 +23,8 @@ el mismo fea_translate_helper.php sin tocar su lógica, solo invertido
|
|||||||
Uso:
|
Uso:
|
||||||
python3 sync_carta_from_prod.py --carta 54495 --dry-run
|
python3 sync_carta_from_prod.py --carta 54495 --dry-run
|
||||||
python3 sync_carta_from_prod.py --carta 54495
|
python3 sync_carta_from_prod.py --carta 54495
|
||||||
|
python3 sync_carta_from_prod.py --ids 54875,54902,54903 --dry-run
|
||||||
|
python3 sync_carta_from_prod.py --ids 54875,54902,54903
|
||||||
|
|
||||||
Tras esto, el resto del ciclo ya existente no cambia:
|
Tras esto, el resto del ciclo ya existente no cambia:
|
||||||
translate_post.py --carta 54495 --langs en,fr,it,pt --status draft
|
translate_post.py --carta 54495 --langs en,fr,it,pt --status draft
|
||||||
@@ -34,8 +36,10 @@ Tras esto, el resto del ciclo ya existente no cambia:
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import argparse
|
import argparse
|
||||||
|
import base64
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
|
import shlex
|
||||||
import subprocess
|
import subprocess
|
||||||
import time
|
import time
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
@@ -44,7 +48,11 @@ WP_CONTAINER = os.environ.get("FEA_WP_CONTAINER", "wordpress-web")
|
|||||||
|
|
||||||
PROD_HOST = os.environ.get("FEA_PROD_SSH_HOST", "")
|
PROD_HOST = os.environ.get("FEA_PROD_SSH_HOST", "")
|
||||||
PROD_PASS = os.environ.get("FEA_PROD_SSH_PASS", "")
|
PROD_PASS = os.environ.get("FEA_PROD_SSH_PASS", "")
|
||||||
PROD_WPLOAD = os.environ.get("FEA_PROD_WPLOAD", "/web/wp-load.php")
|
PROD_WPLOAD = os.environ.get("FEA_PROD_WPLOAD", "/var/www/html/wp-load.php")
|
||||||
|
# Desde el cutover a Hetzner, WordPress vive dentro de Coolify/Docker. Si se
|
||||||
|
# define, el helper se ejecuta en memoria dentro del contenedor; no se escribe
|
||||||
|
# ningún fichero temporal en prod.
|
||||||
|
PROD_DOCKER_CONTAINER = os.environ.get("FEA_PROD_DOCKER_CONTAINER", "")
|
||||||
PROD_HELPER = "/tmp/fea_translate_helper.php"
|
PROD_HELPER = "/tmp/fea_translate_helper.php"
|
||||||
|
|
||||||
HELPER_SRC = Path(__file__).resolve().parent / "fea_translate_helper.php"
|
HELPER_SRC = Path(__file__).resolve().parent / "fea_translate_helper.php"
|
||||||
@@ -74,22 +82,35 @@ def sh(cmd: list[str], *, stdin: str | None = None, timeout: int = 120) -> str:
|
|||||||
|
|
||||||
|
|
||||||
# ── Prod (origen, solo lectura) ─────────────────────────────────────────────
|
# ── Prod (origen, solo lectura) ─────────────────────────────────────────────
|
||||||
_prod_ready = False
|
|
||||||
|
|
||||||
|
|
||||||
def _ssh(remote_cmd: str, *, stdin: str | None = None, timeout: int = 120) -> str:
|
def _ssh(remote_cmd: str, *, stdin: str | None = None, timeout: int = 120) -> str:
|
||||||
|
# El Hetzner nuevo usa auth por clave (ed25519 claude-code@feadulta); sshpass
|
||||||
|
# solo se antepone si hay contraseña configurada (servidor viejo CDMON).
|
||||||
|
if PROD_PASS:
|
||||||
cmd = ["sshpass", "-p", PROD_PASS, "ssh", "-o", "StrictHostKeyChecking=accept-new",
|
cmd = ["sshpass", "-p", PROD_PASS, "ssh", "-o", "StrictHostKeyChecking=accept-new",
|
||||||
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
||||||
|
else:
|
||||||
|
cmd = ["ssh", "-o", "StrictHostKeyChecking=accept-new",
|
||||||
|
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
||||||
return sh(cmd, stdin=stdin, timeout=timeout)
|
return sh(cmd, stdin=stdin, timeout=timeout)
|
||||||
|
|
||||||
|
|
||||||
def prod_helper(subcmd: str, *args: str, stdin: str | None = None) -> str:
|
def prod_helper(subcmd: str, *args: str) -> str:
|
||||||
global _prod_ready
|
"""Run the helper in prod memory; never upload a temporary file to prod."""
|
||||||
if not _prod_ready:
|
helper_php = HELPER_SRC.read_text(encoding="utf-8").replace("<?php", "", 1)
|
||||||
_ssh(f"cat > {PROD_HELPER}", stdin=HELPER_SRC.read_text(encoding="utf-8"))
|
encoded = base64.b64encode(helper_php.encode("utf-8")).decode("ascii")
|
||||||
_prod_ready = True
|
code = f"eval(base64_decode('{encoded}'));"
|
||||||
inner = f"FEA_WP_LOAD={PROD_WPLOAD} php {PROD_HELPER} {subcmd} " + " ".join(args)
|
if PROD_DOCKER_CONTAINER:
|
||||||
return _ssh(inner, stdin=stdin, timeout=180)
|
remote = (
|
||||||
|
f"docker exec -i -e FEA_WP_LOAD={shlex.quote(PROD_WPLOAD)} "
|
||||||
|
f"{shlex.quote(PROD_DOCKER_CONTAINER)} php -r {shlex.quote(code)} -- "
|
||||||
|
+ " ".join(shlex.quote(part) for part in (subcmd, *args))
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
remote = (
|
||||||
|
f"FEA_WP_LOAD={shlex.quote(PROD_WPLOAD)} php -r {shlex.quote(code)} -- "
|
||||||
|
+ " ".join(shlex.quote(part) for part in (subcmd, *args))
|
||||||
|
)
|
||||||
|
return _ssh(remote, timeout=180)
|
||||||
|
|
||||||
|
|
||||||
def prod_read_full(post_id: int) -> dict:
|
def prod_read_full(post_id: int) -> dict:
|
||||||
@@ -110,6 +131,17 @@ def seed_ids_from_carta(carta_id: int) -> list[int]:
|
|||||||
return sorted(ids)
|
return sorted(ids)
|
||||||
|
|
||||||
|
|
||||||
|
def seed_ids_from_csv(raw_ids: str) -> list[int]:
|
||||||
|
"""Parse a deliberate prod→local subset for Mixbot --parte1 (#181)."""
|
||||||
|
try:
|
||||||
|
ids = sorted({int(part.strip()) for part in raw_ids.split(",") if part.strip()})
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError("--ids debe ser una lista CSV de IDs numéricos") from exc
|
||||||
|
if not ids or any(pid <= 0 for pid in ids):
|
||||||
|
raise ValueError("--ids debe contener al menos un ID positivo")
|
||||||
|
return ids
|
||||||
|
|
||||||
|
|
||||||
def collect_related_posts(seed_ids: list[int]) -> tuple[dict[int, dict], list[dict[str, int]]]:
|
def collect_related_posts(seed_ids: list[int]) -> tuple[dict[int, dict], list[dict[str, int]]]:
|
||||||
posts: dict[int, dict] = {}
|
posts: dict[int, dict] = {}
|
||||||
groups: dict[tuple[tuple[str, int], ...], dict[str, int]] = {}
|
groups: dict[tuple[tuple[str, int], ...], dict[str, int]] = {}
|
||||||
@@ -153,7 +185,11 @@ def local_read_safe(post_id: int) -> dict | None:
|
|||||||
return None
|
return None
|
||||||
|
|
||||||
|
|
||||||
def local_clone(post: dict) -> int:
|
def local_clone(post: dict, *, preserve_id: bool = True) -> int:
|
||||||
|
# Conserva el ID remoto como trazabilidad incluso cuando haya que crear un
|
||||||
|
# ID local nuevo por una colisión entre entornos.
|
||||||
|
meta = dict(post.get("meta", {}))
|
||||||
|
meta["fea_prod_source_id"] = [str(post["id"])]
|
||||||
payload = {
|
payload = {
|
||||||
"title": post["title"],
|
"title": post["title"],
|
||||||
"content": post.get("content", ""),
|
"content": post.get("content", ""),
|
||||||
@@ -166,10 +202,14 @@ def local_clone(post: dict) -> int:
|
|||||||
"status": post.get("status"),
|
"status": post.get("status"),
|
||||||
"cats": post.get("cats", []),
|
"cats": post.get("cats", []),
|
||||||
"cat_slugs": post.get("cat_slugs", []),
|
"cat_slugs": post.get("cat_slugs", []),
|
||||||
"meta": post.get("meta", {}),
|
"meta": meta,
|
||||||
}
|
}
|
||||||
|
if preserve_id:
|
||||||
out = local_helper("clone", str(post["id"]), post.get("lang") or "es", STATUS,
|
out = local_helper("clone", str(post["id"]), post.get("lang") or "es", STATUS,
|
||||||
stdin=json.dumps(payload)).strip()
|
stdin=json.dumps(payload)).strip()
|
||||||
|
else:
|
||||||
|
out = local_helper("clone_new", post.get("lang") or "es", STATUS,
|
||||||
|
stdin=json.dumps(payload)).strip()
|
||||||
return int(out)
|
return int(out)
|
||||||
|
|
||||||
|
|
||||||
@@ -179,13 +219,9 @@ def local_save_group(group: dict[str, int]) -> dict[str, int]:
|
|||||||
|
|
||||||
|
|
||||||
# ── Main ─────────────────────────────────────────────────────────────────────
|
# ── Main ─────────────────────────────────────────────────────────────────────
|
||||||
def sync_carta(carta_id: int, *, dry_run: bool) -> int:
|
def sync_posts(seed_ids: list[int], *, dry_run: bool, source_label: str,
|
||||||
seed_ids = seed_ids_from_carta(carta_id)
|
remap_conflicts: bool = False) -> int:
|
||||||
log(f"Cluster descubierto desde fea_parse_carta_sections(prod#{carta_id}): "
|
log(f"Lote {source_label}: {len(seed_ids)} post(s) -> {seed_ids}")
|
||||||
f"{len(seed_ids)} post(s) -> {seed_ids}")
|
|
||||||
if not seed_ids or seed_ids == [carta_id]:
|
|
||||||
log(" ⚠️ El parser no resolvió ningún artículo enlazado (¿carta sin publicar aún, "
|
|
||||||
"o secciones sin encabezados reconocibles?). Revisa antes de continuar.")
|
|
||||||
|
|
||||||
posts, groups = collect_related_posts(seed_ids)
|
posts, groups = collect_related_posts(seed_ids)
|
||||||
|
|
||||||
@@ -196,51 +232,82 @@ def sync_carta(carta_id: int, *, dry_run: bool) -> int:
|
|||||||
conflicts.append((pid, existing.get("title", ""), p.get("title", "")))
|
conflicts.append((pid, existing.get("title", ""), p.get("title", "")))
|
||||||
|
|
||||||
if conflicts:
|
if conflicts:
|
||||||
log(f" ⚠️ {len(conflicts)} CONFLICTO(S): el ID ya existe en local con OTRO contenido "
|
log(f" ⚠️ {len(conflicts)} CONFLICTO(S): el ID ya existe en local con OTRO contenido:")
|
||||||
f"y va a ser SOBRESCRITO:")
|
|
||||||
for pid, old_title, new_title in conflicts:
|
for pid, old_title, new_title in conflicts:
|
||||||
log(f" #{pid}: local actual «{old_title[:50]}» -> prod «{new_title[:50]}»")
|
log(f" #{pid}: local actual «{old_title[:50]}» | prod «{new_title[:50]}»")
|
||||||
|
if not dry_run and not remap_conflicts:
|
||||||
|
raise RuntimeError(
|
||||||
|
"Importación cancelada: usar --remap-conflicts para crear IDs locales nuevos; "
|
||||||
|
"nunca se sobrescriben posts locales por un choque de IDs entre entornos."
|
||||||
|
)
|
||||||
|
|
||||||
if dry_run:
|
if dry_run:
|
||||||
for pid, p in posts.items():
|
for pid, p in posts.items():
|
||||||
log(f" PULL prod#{pid} [{p.get('lang','?')}] status={p.get('status')} "
|
action = "REMAPPING to new local ID" if pid in {c[0] for c in conflicts} else "PULL"
|
||||||
|
log(f" {action} prod#{pid} [{p.get('lang','?')}] status={p.get('status')} "
|
||||||
f"slug={p.get('slug','')} «{p.get('title','')[:50]}»")
|
f"slug={p.get('slug','')} «{p.get('title','')[:50]}»")
|
||||||
for group in groups:
|
for group in groups:
|
||||||
log(f" GROUP {group}")
|
log(f" GROUP {group}")
|
||||||
log("DRY-RUN: nada escrito en local.")
|
log("DRY-RUN: nada escrito en local.")
|
||||||
return 0
|
return 0
|
||||||
|
|
||||||
|
conflict_ids = {pid for pid, _, _ in conflicts}
|
||||||
|
id_map: dict[int, int] = {}
|
||||||
for pid, p in posts.items():
|
for pid, p in posts.items():
|
||||||
new_id = local_clone(p)
|
new_id = local_clone(p, preserve_id=pid not in conflict_ids)
|
||||||
|
id_map[pid] = new_id
|
||||||
if new_id != pid:
|
if new_id != pid:
|
||||||
log(f" ⚠️ prod#{pid} se clonó como local#{new_id} — ID NO preservado, revisar a mano.")
|
log(f" remap prod#{pid} -> local#{new_id} [{p.get('lang','?')}] «{p['title'][:45]}»")
|
||||||
else:
|
else:
|
||||||
log(f" clone prod#{pid} -> local#{new_id} [{p.get('lang','?')}] «{p['title'][:45]}»")
|
log(f" clone prod#{pid} -> local#{new_id} [{p.get('lang','?')}] «{p['title'][:45]}»")
|
||||||
|
|
||||||
for group in groups:
|
for group in groups:
|
||||||
if len(group) < 2:
|
local_group = {lang: id_map.get(pid, pid) for lang, pid in group.items()}
|
||||||
|
if len(local_group) < 2:
|
||||||
continue
|
continue
|
||||||
saved = local_save_group(group)
|
saved = local_save_group(local_group)
|
||||||
log(f" group enlazado en local {saved}")
|
log(f" group enlazado en local {saved}")
|
||||||
|
|
||||||
log(f"FIN sync prod->local. carta={carta_id} posts={len(posts)} conflictos_previos={len(conflicts)}")
|
log(f"FIN sync prod->local. fuente={source_label} posts={len(posts)} conflictos_previos={len(conflicts)}")
|
||||||
return 0
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
def sync_carta(carta_id: int, *, dry_run: bool, remap_conflicts: bool = False) -> int:
|
||||||
|
seed_ids = seed_ids_from_carta(carta_id)
|
||||||
|
if not seed_ids or seed_ids == [carta_id]:
|
||||||
|
log(" ⚠️ El parser no resolvió ningún artículo enlazado (¿carta sin publicar aún, "
|
||||||
|
"o secciones sin encabezados reconocibles?). Revisa antes de continuar.")
|
||||||
|
return sync_posts(seed_ids, dry_run=dry_run, source_label=f"carta prod#{carta_id}",
|
||||||
|
remap_conflicts=remap_conflicts)
|
||||||
|
|
||||||
|
|
||||||
|
def sync_ids(raw_ids: str, *, dry_run: bool, remap_conflicts: bool = False) -> int:
|
||||||
|
return sync_posts(seed_ids_from_csv(raw_ids), dry_run=dry_run,
|
||||||
|
source_label="IDs explícitos de Mixbot --parte1",
|
||||||
|
remap_conflicts=remap_conflicts)
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
def main() -> int:
|
||||||
ap = argparse.ArgumentParser(
|
ap = argparse.ArgumentParser(
|
||||||
description="Copia una carta (y su cluster de artículos) de PROD a LOCAL preservando IDs.")
|
description="Copia una carta (y su cluster de artículos) de PROD a LOCAL preservando IDs.")
|
||||||
ap.add_argument("--carta", type=int, required=True, help="ID del post ES de la carta en PROD.")
|
group = ap.add_mutually_exclusive_group(required=True)
|
||||||
|
group.add_argument("--carta", type=int, help="ID del post ES de la carta en PROD.")
|
||||||
|
group.add_argument("--ids", help="Lista CSV explícita de posts ES para Mixbot --parte1 (#181).")
|
||||||
ap.add_argument("--dry-run", action="store_true", help="Solo muestra el plan; no escribe en local.")
|
ap.add_argument("--dry-run", action="store_true", help="Solo muestra el plan; no escribe en local.")
|
||||||
|
ap.add_argument("--remap-conflicts", action="store_true",
|
||||||
|
help="Ante IDs ocupados localmente, crea IDs nuevos y guarda fea_prod_source_id; nunca sobrescribe.")
|
||||||
args = ap.parse_args()
|
args = ap.parse_args()
|
||||||
|
|
||||||
if not PROD_HOST or not PROD_PASS:
|
if not PROD_HOST:
|
||||||
raise SystemExit(
|
raise SystemExit(
|
||||||
"Faltan FEA_PROD_SSH_HOST / FEA_PROD_SSH_PASS en el entorno.\n"
|
"Falta FEA_PROD_SSH_HOST en el entorno (FEA_PROD_SSH_PASS es opcional: "
|
||||||
|
"el Hetzner nuevo usa auth por clave).\n"
|
||||||
"Antes de ejecutar: source ~/.hermes/profiles/feadulta/.env"
|
"Antes de ejecutar: source ~/.hermes/profiles/feadulta/.env"
|
||||||
)
|
)
|
||||||
|
|
||||||
return sync_carta(args.carta, dry_run=args.dry_run)
|
if args.carta:
|
||||||
|
return sync_carta(args.carta, dry_run=args.dry_run, remap_conflicts=args.remap_conflicts)
|
||||||
|
return sync_ids(args.ids, dry_run=args.dry_run, remap_conflicts=args.remap_conflicts)
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|||||||
@@ -16,8 +16,10 @@ coincidencia local↔prod cuando prod va por detrás.
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import argparse
|
import argparse
|
||||||
|
import base64
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
|
import shlex
|
||||||
import subprocess
|
import subprocess
|
||||||
import time
|
import time
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
@@ -29,12 +31,20 @@ DB_NAME = os.environ.get("FEA_DB_NAME", "wordpress_db")
|
|||||||
DB_USER = os.environ.get("FEA_DB_USER", "wordpress_user")
|
DB_USER = os.environ.get("FEA_DB_USER", "wordpress_user")
|
||||||
DB_PASS = os.environ.get("FEA_DB_PASS", "wordpress_pass")
|
DB_PASS = os.environ.get("FEA_DB_PASS", "wordpress_pass")
|
||||||
|
|
||||||
PROD_HOST = os.environ.get("FEA_PROD_HOST", "feadulta@134.0.10.170")
|
PROD_HOST = os.environ.get("FEA_PROD_SSH_HOST", "")
|
||||||
PROD_PASS = os.environ.get("FEA_PROD_PASS", "C6c2A!mAl3Wj.BQF")
|
PROD_PASS = os.environ.get("FEA_PROD_SSH_PASS", "")
|
||||||
PROD_WPLOAD = os.environ.get("FEA_PROD_WPLOAD", "/web/wp-load.php")
|
PROD_WPLOAD = os.environ.get("FEA_PROD_WPLOAD", "/var/www/html/wp-load.php")
|
||||||
PROD_HELPER = "/tmp/fea_translate_helper.php"
|
# Desde el cutover a Hetzner, WordPress vive dentro de Coolify/Docker. Si se
|
||||||
|
# define, el helper se ejecuta en memoria dentro del contenedor; no se escribe
|
||||||
|
# ningún fichero temporal en prod. Espejo de sync_carta_from_prod.py.
|
||||||
|
PROD_DOCKER_CONTAINER = os.environ.get("FEA_PROD_DOCKER_CONTAINER", "")
|
||||||
|
|
||||||
HELPER_SRC = Path(__file__).resolve().parent / "fea_translate_helper.php"
|
HELPER_SRC = Path(__file__).resolve().parent / "fea_translate_helper.php"
|
||||||
|
# El helper se ejecuta en memoria con `php -r` en prod. Así no se copia ningún
|
||||||
|
# fichero a /tmp remoto y stdin queda disponible para el payload JSON.
|
||||||
|
HELPER_EVAL_CODE = "eval(base64_decode(" + repr(
|
||||||
|
base64.b64encode(HELPER_SRC.read_text(encoding="utf-8").removeprefix("<?php").encode("utf-8")).decode("ascii")
|
||||||
|
) + "));"
|
||||||
LOCAL_HELPER_DST = "/tmp/fea_translate_helper.php"
|
LOCAL_HELPER_DST = "/tmp/fea_translate_helper.php"
|
||||||
STATE_FILE = Path(os.environ.get("FEA_SYNC_STATE", "/tmp/feadulta-sync-state.json"))
|
STATE_FILE = Path(os.environ.get("FEA_SYNC_STATE", "/tmp/feadulta-sync-state.json"))
|
||||||
LOG_FILE = Path(os.environ.get("FEA_SYNC_LOG", "/tmp/feadulta-sync.log"))
|
LOG_FILE = Path(os.environ.get("FEA_SYNC_LOG", "/tmp/feadulta-sync.log"))
|
||||||
@@ -107,8 +117,20 @@ def local_read_full(post_id: int) -> dict:
|
|||||||
|
|
||||||
|
|
||||||
def local_translation_pairs() -> list[tuple[int, int]]:
|
def local_translation_pairs() -> list[tuple[int, int]]:
|
||||||
q = ("SELECT post_id, meta_value FROM wp_postmeta "
|
"""Devuelve (traducción_local, origen_ES_en_prod).
|
||||||
"WHERE meta_key='traduccion_origen' ORDER BY CAST(meta_value AS UNSIGNED), post_id;")
|
|
||||||
|
Normalmente ambos entornos compartían IDs. Para importaciones remapeadas,
|
||||||
|
la carta ES local guarda `fea_prod_source_id`; esa referencia remota tiene
|
||||||
|
prioridad y evita enlazar una traducción con un post distinto de prod.
|
||||||
|
"""
|
||||||
|
q = (
|
||||||
|
"SELECT tr.post_id, COALESCE(NULLIF(src.meta_value,''), tr.meta_value) "
|
||||||
|
"FROM wp_postmeta tr "
|
||||||
|
"LEFT JOIN wp_postmeta src ON src.post_id=CAST(tr.meta_value AS UNSIGNED) "
|
||||||
|
"AND src.meta_key='fea_prod_source_id' "
|
||||||
|
"WHERE tr.meta_key='traduccion_origen' "
|
||||||
|
"ORDER BY CAST(COALESCE(NULLIF(src.meta_value,''), tr.meta_value) AS UNSIGNED), tr.post_id;"
|
||||||
|
)
|
||||||
out = sh(["docker", "exec", DB_CONTAINER, "mysql", f"-u{DB_USER}", f"-p{DB_PASS}",
|
out = sh(["docker", "exec", DB_CONTAINER, "mysql", f"-u{DB_USER}", f"-p{DB_PASS}",
|
||||||
DB_NAME, "-N", "-e", q])
|
DB_NAME, "-N", "-e", q])
|
||||||
pairs = []
|
pairs = []
|
||||||
@@ -155,21 +177,31 @@ def collect_related_posts(seed_ids: list[int]) -> tuple[dict[int, dict], list[di
|
|||||||
|
|
||||||
|
|
||||||
# ── Prod ─────────────────────────────────────────────────────────────────────
|
# ── Prod ─────────────────────────────────────────────────────────────────────
|
||||||
_prod_ready = False
|
|
||||||
|
|
||||||
|
|
||||||
def _ssh(remote_cmd: str, *, stdin: str | None = None, timeout: int = 120) -> str:
|
def _ssh(remote_cmd: str, *, stdin: str | None = None, timeout: int = 120) -> str:
|
||||||
|
# El Hetzner nuevo usa auth por clave (ed25519 claude-code@feadulta); sshpass
|
||||||
|
# solo se antepone si hay contraseña configurada (servidor viejo CDMON).
|
||||||
|
if PROD_PASS:
|
||||||
cmd = ["sshpass", "-p", PROD_PASS, "ssh", "-o", "StrictHostKeyChecking=accept-new",
|
cmd = ["sshpass", "-p", PROD_PASS, "ssh", "-o", "StrictHostKeyChecking=accept-new",
|
||||||
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
||||||
|
else:
|
||||||
|
cmd = ["ssh", "-o", "StrictHostKeyChecking=accept-new",
|
||||||
|
"-o", "ConnectTimeout=20", PROD_HOST, remote_cmd]
|
||||||
return sh(cmd, stdin=stdin, timeout=timeout)
|
return sh(cmd, stdin=stdin, timeout=timeout)
|
||||||
|
|
||||||
|
|
||||||
def prod_helper(subcmd: str, *args: str, stdin: str | None = None) -> str:
|
def prod_helper(subcmd: str, *args: str, stdin: str | None = None) -> str:
|
||||||
global _prod_ready
|
"""Ejecuta el helper PHP en memoria; no crea ficheros temporales en prod."""
|
||||||
if not _prod_ready:
|
argv = " ".join(shlex.quote(x) for x in (subcmd, *args))
|
||||||
_ssh(f"cat > {PROD_HELPER}", stdin=HELPER_SRC.read_text(encoding="utf-8"))
|
if PROD_DOCKER_CONTAINER:
|
||||||
_prod_ready = True
|
inner = (
|
||||||
inner = f"FEA_WP_LOAD={PROD_WPLOAD} php {PROD_HELPER} {subcmd} " + " ".join(args)
|
f"docker exec -i -e FEA_WP_LOAD={shlex.quote(PROD_WPLOAD)} "
|
||||||
|
f"{shlex.quote(PROD_DOCKER_CONTAINER)} php -r {shlex.quote(HELPER_EVAL_CODE)} {argv}"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
inner = (
|
||||||
|
f"FEA_WP_LOAD={shlex.quote(PROD_WPLOAD)} "
|
||||||
|
f"php -r {shlex.quote(HELPER_EVAL_CODE)} {argv}"
|
||||||
|
)
|
||||||
return _ssh(inner, stdin=stdin, timeout=180)
|
return _ssh(inner, stdin=stdin, timeout=180)
|
||||||
|
|
||||||
|
|
||||||
@@ -256,7 +288,12 @@ def deploy_fixed_ids(seed_ids: list[int], *, keep_existing: set[int], dry_run: b
|
|||||||
|
|
||||||
|
|
||||||
# ── Main legado ──────────────────────────────────────────────────────────────
|
# ── Main legado ──────────────────────────────────────────────────────────────
|
||||||
def legacy_sync(limit: int, origin: int) -> int:
|
def legacy_sync(limit: int, origin: int, *, dry_run: bool = False) -> int:
|
||||||
|
"""Sincroniza por origen ES+idioma, sin reutilizar IDs locales en prod.
|
||||||
|
|
||||||
|
Este es el modo seguro si producción ha avanzado y sus IDs ya pueden
|
||||||
|
colisionar con traducciones creadas en el entorno local.
|
||||||
|
"""
|
||||||
state = load_state()
|
state = load_state()
|
||||||
pairs = local_translation_pairs()
|
pairs = local_translation_pairs()
|
||||||
if origin:
|
if origin:
|
||||||
@@ -280,6 +317,10 @@ def legacy_sync(limit: int, origin: int) -> int:
|
|||||||
if key in state["done"]:
|
if key in state["done"]:
|
||||||
n_skip += 1
|
n_skip += 1
|
||||||
continue
|
continue
|
||||||
|
if dry_run:
|
||||||
|
log(f" PLAN {key}: crearía traducción [{lang}] «{t['title'][:45]}»")
|
||||||
|
n_ok += 1
|
||||||
|
continue
|
||||||
try:
|
try:
|
||||||
new_id = prod_create(src_origin, lang, t["title"], t["content"])
|
new_id = prod_create(src_origin, lang, t["title"], t["content"])
|
||||||
state["done"][key] = new_id
|
state["done"][key] = new_id
|
||||||
@@ -292,8 +333,10 @@ def legacy_sync(limit: int, origin: int) -> int:
|
|||||||
n_err += 1
|
n_err += 1
|
||||||
log(f" {key} ERROR: {exc}")
|
log(f" {key} ERROR: {exc}")
|
||||||
|
|
||||||
|
if not dry_run:
|
||||||
save_state(state)
|
save_state(state)
|
||||||
log(f"FIN sync legado. nuevos={n_ok} saltados={n_skip} errores={n_err}. Estado: {STATE_FILE}")
|
mode = "DRY-RUN" if dry_run else "SYNC"
|
||||||
|
log(f"FIN sync legado {mode}. nuevos/plan={n_ok} saltados={n_skip} errores={n_err}. Estado: {STATE_FILE}")
|
||||||
log("Recuerda en prod: ejecutar remap_translation_cats.php si alguna quedó sin categoría traducida.")
|
log("Recuerda en prod: ejecutar remap_translation_cats.php si alguna quedó sin categoría traducida.")
|
||||||
return 0
|
return 0
|
||||||
|
|
||||||
@@ -318,7 +361,7 @@ def main() -> int:
|
|||||||
keep_existing = set(parse_csv_ints(args.keep_existing))
|
keep_existing = set(parse_csv_ints(args.keep_existing))
|
||||||
return deploy_fixed_ids(seed_ids, keep_existing=keep_existing, dry_run=args.dry_run)
|
return deploy_fixed_ids(seed_ids, keep_existing=keep_existing, dry_run=args.dry_run)
|
||||||
|
|
||||||
return legacy_sync(args.limit, args.origin)
|
return legacy_sync(args.limit, args.origin, dry_run=args.dry_run)
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|||||||
@@ -271,7 +271,8 @@ def save_state(state: dict) -> None:
|
|||||||
|
|
||||||
|
|
||||||
# ── Orquestación ─────────────────────────────────────────────────────────────
|
# ── Orquestación ─────────────────────────────────────────────────────────────
|
||||||
def process_post(post_id: int, langs: list[str], status: str, force: bool, state: dict) -> None:
|
def process_post(post_id: int, langs: list[str], status: str, force: bool, state: dict,
|
||||||
|
*, dry_run: bool = False) -> None:
|
||||||
src = read_post(post_id)
|
src = read_post(post_id)
|
||||||
if src.get("lang") and src["lang"] != "es":
|
if src.get("lang") and src["lang"] != "es":
|
||||||
log(f"#{post_id} no es ES (lang={src['lang']}) — saltado")
|
log(f"#{post_id} no es ES (lang={src['lang']}) — saltado")
|
||||||
@@ -286,8 +287,14 @@ def process_post(post_id: int, langs: list[str], status: str, force: bool, state
|
|||||||
state["done"][key] = existing
|
state["done"][key] = existing
|
||||||
continue
|
continue
|
||||||
if existing and force:
|
if existing and force:
|
||||||
|
if dry_run:
|
||||||
|
log(f" {lang}: PLAN regenerar traducción previa #{existing}")
|
||||||
|
continue
|
||||||
php_helper("unlink", str(post_id), lang)
|
php_helper("unlink", str(post_id), lang)
|
||||||
log(f" {lang}: --force, eliminada traducción previa #{existing}")
|
log(f" {lang}: --force, eliminada traducción previa #{existing}")
|
||||||
|
if dry_run:
|
||||||
|
log(f" {lang}: PLAN traducir con {ENGINE} y crear como {status}")
|
||||||
|
continue
|
||||||
try:
|
try:
|
||||||
t0 = time.time()
|
t0 = time.time()
|
||||||
title = translate_text(src["title"], lang, is_title=True)
|
title = translate_text(src["title"], lang, is_title=True)
|
||||||
@@ -312,6 +319,7 @@ def main() -> int:
|
|||||||
ap.add_argument("--langs", default="en,fr,it,pt", help="Idiomas destino separados por coma.")
|
ap.add_argument("--langs", default="en,fr,it,pt", help="Idiomas destino separados por coma.")
|
||||||
ap.add_argument("--status", default="draft", choices=["draft", "publish"], help="Estado de la traducción.")
|
ap.add_argument("--status", default="draft", choices=["draft", "publish"], help="Estado de la traducción.")
|
||||||
ap.add_argument("--force", action="store_true", help="Regenera aunque ya exista la traducción.")
|
ap.add_argument("--force", action="store_true", help="Regenera aunque ya exista la traducción.")
|
||||||
|
ap.add_argument("--dry-run", action="store_true", help="Muestra el plan sin llamar al LLM ni escribir en WordPress/estado.")
|
||||||
args = ap.parse_args()
|
args = ap.parse_args()
|
||||||
|
|
||||||
langs = [l.strip() for l in args.langs.split(",") if l.strip() in LANG_NAMES]
|
langs = [l.strip() for l in args.langs.split(",") if l.strip() in LANG_NAMES]
|
||||||
@@ -329,9 +337,11 @@ def main() -> int:
|
|||||||
|
|
||||||
state = load_state()
|
state = load_state()
|
||||||
for pid in ids:
|
for pid in ids:
|
||||||
process_post(pid, langs, args.status, args.force, state)
|
process_post(pid, langs, args.status, args.force, state, dry_run=args.dry_run)
|
||||||
|
if not args.dry_run:
|
||||||
save_state(state)
|
save_state(state)
|
||||||
log(f"FIN. {len(state['done'])} traducciones registradas, "
|
mode = "DRY-RUN" if args.dry_run else "FIN"
|
||||||
|
log(f"{mode}. {len(state['done'])} traducciones registradas, "
|
||||||
f"{len(state.get('errors', {}))} errores. Estado: {STATE_FILE}")
|
f"{len(state.get('errors', {}))} errores. Estado: {STATE_FILE}")
|
||||||
return 0
|
return 0
|
||||||
|
|
||||||
|
|||||||
Executable
+142
@@ -0,0 +1,142 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Cron del backlog de TTS por autor (issue rafa/feadulta#188).
|
||||||
|
# Corre cada 5 h los lunes, viernes, sábados y domingos, aprovechando la cuota
|
||||||
|
# ociosa de MiniMax para locutar artículos antiguos. Por ventana:
|
||||||
|
# 1) Mide la cuota y CALCULA el tamaño de la tanda para llenar la ventana hasta
|
||||||
|
# el objetivo. Un tamaño fijo desaprovecha: deja la ventana a medias cuando
|
||||||
|
# está libre, y no cabe cuando está medio usada. Al dimensionar por hueco
|
||||||
|
# libre, además, deja de importar dónde caiga el cron respecto a la ventana.
|
||||||
|
# 2) tts_produce.py --autor ... --max N: la cola sale de la BD, así que esto es
|
||||||
|
# idempotente por construcción — lo ya locutado no vuelve a salir.
|
||||||
|
# SOLO LOCAL: no toca producción. Publicar en prod es sync_audio_to_prod.py, que
|
||||||
|
# está bloqueado hasta después del cutover a Hetzner (#180).
|
||||||
|
# flock evita solapes si una ventana se alargara. Log por día.
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
export PATH="/home/rafa/.local/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin"
|
||||||
|
export HOME="/home/rafa"
|
||||||
|
|
||||||
|
# El repo NO se fija a mano: se deduce de dónde vive este script. El checkout
|
||||||
|
# principal cambia de rama a menudo, y en las ramas que no llevan estos scripts
|
||||||
|
# el cron se quedaba llamando a un fichero inexistente y fallaba en silencio
|
||||||
|
# (pasó del 5 al 8 de agosto de 2026: 7 ventanas perdidas). Por eso el cron
|
||||||
|
# apunta al worktree fijo ~/worktrees/fea-tts-backlog, y desde aquí se
|
||||||
|
# autodetecta: donde esté el script, ahí está su repo.
|
||||||
|
REPO="${FEA_TTS_REPO:-$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)}"
|
||||||
|
PY="/home/rafa/tts-local/xtts-venv/bin/python"
|
||||||
|
QUOTA="/home/rafa/ytsummaries/scripts/quota.py"
|
||||||
|
WORK="/tmp/fea-tts-backlog"
|
||||||
|
LOG="$WORK/cron-$(date +%F).log"
|
||||||
|
LOCK="$WORK/cron.lock"
|
||||||
|
|
||||||
|
# Cola: cambiar aquí (o por entorno) para pasar de un autor a otro.
|
||||||
|
# 382 Fray Marcos · 383 Pagola · 774 Sicre · 386 Arregi
|
||||||
|
AUTOR="${FEA_TTS_AUTOR:-382}"
|
||||||
|
DESDE="${FEA_TTS_DESDE:-2025}"
|
||||||
|
HASTA="${FEA_TTS_HASTA:-2026}"
|
||||||
|
|
||||||
|
# Coste medido de un audio, en DÉCIMAS de punto porcentual (aritmética entera en
|
||||||
|
# bash). Dos tandas de 10 el 2026-08-02: la ventana de 5 h fue 0→45→88 (~4,4 pts
|
||||||
|
# por audio) y la semanal 10→14→18 (~0,4 pts). Artículos de Fray Marcos de
|
||||||
|
# 4.000-5.000 caracteres; si se locuta a otro autor con textos mucho más largos,
|
||||||
|
# revisar estos números con un par de tandas.
|
||||||
|
COSTE_5H="${FEA_TTS_COSTE_5H:-44}"
|
||||||
|
COSTE_SEM="${FEA_TTS_COSTE_SEM:-4}"
|
||||||
|
|
||||||
|
# Objetivos de llenado (%). Dejar cuota sin usar al llegar el reset es tirarla.
|
||||||
|
OBJ_5H="${FEA_TTS_OBJ_5H:-90}"
|
||||||
|
OBJ_SEM="${FEA_TTS_OBJ_SEM:-85}"
|
||||||
|
|
||||||
|
# La semanal NO se gasta a tope en cada ventana: se REPARTE entre las ventanas
|
||||||
|
# que quedan hasta su reset. Llenar cada ventana de 5 h al 90 % son ~8 puntos de
|
||||||
|
# semanal, y hay 20 ventanas activas por semana: 160 puntos para un presupuesto
|
||||||
|
# de 85. Sin reparto, domingo y lunes se lo comen y el fin de semana se queda a
|
||||||
|
# cero. Con reparto sale ~10 audios por ventana, y en la última ventana de la
|
||||||
|
# semana el reparto vale todo lo que sobre, así que tampoco queda cuota sin usar.
|
||||||
|
|
||||||
|
# Tope de seguridad por tanda y override manual (FEA_TTS_BATCH fija el tamaño y
|
||||||
|
# se salta el cálculo).
|
||||||
|
MAX_BATCH="${FEA_TTS_MAX_BATCH:-25}"
|
||||||
|
|
||||||
|
mkdir -p "$WORK"
|
||||||
|
cd "$REPO" || exit 1
|
||||||
|
|
||||||
|
ts() { date +'%F %T'; }
|
||||||
|
exec 9>"$LOCK"
|
||||||
|
if ! flock -n 9; then
|
||||||
|
echo "[$(ts)] otra corrida en curso, salto." >> "$LOG"; exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "[$(ts)] === cron TTS backlog start (autor=$AUTOR $DESDE-$HASTA) ===" >> "$LOG"
|
||||||
|
|
||||||
|
# 1) Medir la cuota, y contar cuántas ventanas del cron quedan hasta que se
|
||||||
|
# reinicie la semanal — es el denominador del reparto.
|
||||||
|
read -r PCT5 PCTW HSEM VENTANAS <<< "$(python3 "$QUOTA" --json --no-local 2>/dev/null | python3 -c '
|
||||||
|
import json, sys
|
||||||
|
from datetime import datetime, timedelta, timezone
|
||||||
|
|
||||||
|
# DEBEN COINCIDIR CON EL CRONTAB: 0 */5 * * 1,5,6,0
|
||||||
|
HORAS = {0, 5, 10, 15, 20}
|
||||||
|
DIAS = {0, 4, 5, 6} # lun, vie, sab, dom en datetime.weekday()
|
||||||
|
|
||||||
|
def pct(v):
|
||||||
|
# OJO: 0.0 es un valor legítimo (ventana entera libre) y es falsy en Python.
|
||||||
|
# Un `v or 100` aquí aborta la tanda justo cuando hay toda la cuota disponible.
|
||||||
|
return int(v) if v is not None else 100
|
||||||
|
|
||||||
|
try:
|
||||||
|
d = json.load(sys.stdin)
|
||||||
|
m = next(p for p in d["providers"] if p["provider"] == "minimax" and p.get("ok"))
|
||||||
|
horas, ventanas = 999, 1
|
||||||
|
try:
|
||||||
|
fin = datetime.fromisoformat(m["week_reset"]).astimezone()
|
||||||
|
ahora = datetime.now().astimezone()
|
||||||
|
horas = max(int((fin - ahora).total_seconds() // 3600), 0)
|
||||||
|
# Esta corrida cuenta como una; se suman las que quedan programadas.
|
||||||
|
t = (ahora + timedelta(hours=1)).replace(minute=0, second=0, microsecond=0)
|
||||||
|
while t < fin:
|
||||||
|
if t.hour in HORAS and t.weekday() in DIAS:
|
||||||
|
ventanas += 1
|
||||||
|
t += timedelta(hours=1)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
print(pct(m.get("five_h_pct")), pct(m.get("week_pct")), horas, ventanas)
|
||||||
|
except Exception:
|
||||||
|
print(100, 100, 999, 1) # sin lectura fiable de cuota, no se gasta
|
||||||
|
' 2>/dev/null || echo "100 100 999 1")"
|
||||||
|
[ "${VENTANAS:-0}" -lt 1 ] && VENTANAS=1
|
||||||
|
|
||||||
|
# 2) Dimensionar la tanda. Dos límites, manda el más restrictivo:
|
||||||
|
# - la ventana de 5 h: se llena hasta OBJ_5H aquí y ahora;
|
||||||
|
# - la semanal: solo la parte que le toca a esta ventana de lo que queda.
|
||||||
|
CABE_5H=$(( ((OBJ_5H - PCT5) * 10) / COSTE_5H ))
|
||||||
|
CABE_SEM=$(( (((OBJ_SEM - PCTW) * 10) / VENTANAS) / COSTE_SEM ))
|
||||||
|
[ "$CABE_5H" -lt 0 ] && CABE_5H=0
|
||||||
|
[ "$CABE_SEM" -lt 0 ] && CABE_SEM=0
|
||||||
|
|
||||||
|
BATCH=$CABE_5H
|
||||||
|
[ "$CABE_SEM" -lt "$BATCH" ] && BATCH=$CABE_SEM
|
||||||
|
[ "$BATCH" -gt "$MAX_BATCH" ] && BATCH=$MAX_BATCH
|
||||||
|
# Override manual: fija el tamaño y se salta todo el cálculo.
|
||||||
|
[ -n "${FEA_TTS_BATCH:-}" ] && BATCH="$FEA_TTS_BATCH"
|
||||||
|
|
||||||
|
echo "[$(ts)] MiniMax 5h=${PCT5}% semana=${PCTW}% · reset semanal en ${HSEM}h, ${VENTANAS} ventanas por delante" >> "$LOG"
|
||||||
|
echo "[$(ts)] Caben: ${CABE_5H} por la de 5h, ${CABE_SEM} por el reparto semanal → tanda de ${BATCH}" >> "$LOG"
|
||||||
|
|
||||||
|
if [ "$BATCH" -lt 1 ]; then
|
||||||
|
echo "[$(ts)] ABORT: no cabe ni un audio sin pasarse del objetivo; salto esta ventana." >> "$LOG"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 3) Tanda. tts_produce.py ya para solo ante rc 2056/1039 (cuota/rate limit).
|
||||||
|
# Una tanda larga puede desbordar el reset de 5 h (~2,6 min por audio): no pasa
|
||||||
|
# nada, lo que sobra lo absorbe la ventana siguiente y la próxima corrida la
|
||||||
|
# mide y se redimensiona sola. Cortar por tiempo dejaría cuota sin gastar.
|
||||||
|
echo "[$(ts)] tts_produce --autor $AUTOR --desde $DESDE --hasta $HASTA --max $BATCH ..." >> "$LOG"
|
||||||
|
"$PY" scripts/tts_produce.py --autor "$AUTOR" --desde "$DESDE" --hasta "$HASTA" \
|
||||||
|
--max "$BATCH" >> "$LOG" 2>&1
|
||||||
|
|
||||||
|
# 4) Cuánto queda tras la tanda (recuento fresco de la BD, barato y sin cuota).
|
||||||
|
QUEDAN="$("$PY" scripts/tts_produce.py --autor "$AUTOR" --desde "$DESDE" --hasta "$HASTA" \
|
||||||
|
--dry-run 2>/dev/null | sed -n 's/.*Cola: \([0-9]*\) posts.*/\1/p' | tail -1)"
|
||||||
|
echo "[$(ts)] === cron TTS backlog done. Pendientes autor $AUTOR $DESDE-$HASTA: ${QUEDAN:-?} ===" >> "$LOG"
|
||||||
+83
-8
@@ -6,9 +6,18 @@ Reanudable (meta fea_audio_done) y con freno ante la cuota (para tras N fallos
|
|||||||
seguidos). NO toca el front; solo genera el mp3 y asocia la URL al post (meta
|
seguidos). NO toca el front; solo genera el mp3 y asocia la URL al post (meta
|
||||||
fea_audio_url).
|
fea_audio_url).
|
||||||
|
|
||||||
|
Dos modos de cola:
|
||||||
|
- cartas (por defecto): FEA_TTS_CARTAS / --cartas / --ids. Es el flujo de la
|
||||||
|
carta semanal, que tiene prioridad y no cambia.
|
||||||
|
- backlog por autor: --autor 382 [--desde 2025] [--hasta 2026] [--max 15].
|
||||||
|
La cola sale de `listpending` en fea_post_io.php (posts ES publicados sin
|
||||||
|
audio, más recientes primero). Esa consulta ES la idempotencia: no hay
|
||||||
|
fichero de estado, lo ya locutado deja de salir solo.
|
||||||
|
|
||||||
Lanzar: nohup ~/tts-local/xtts-venv/bin/python scripts/tts_produce.py > /tmp/feadulta-tts-prod.out 2>&1 &
|
Lanzar: nohup ~/tts-local/xtts-venv/bin/python scripts/tts_produce.py > /tmp/feadulta-tts-prod.out 2>&1 &
|
||||||
Log: /tmp/feadulta-tts-prod.log
|
Log: /tmp/feadulta-tts-prod.log
|
||||||
"""
|
"""
|
||||||
|
import argparse
|
||||||
import os
|
import os
|
||||||
import shutil
|
import shutil
|
||||||
import subprocess
|
import subprocess
|
||||||
@@ -20,14 +29,15 @@ sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
|||||||
import minimax_tts as mm # get_post_text, add_pauses, t2a, OUT
|
import minimax_tts as mm # get_post_text, add_pauses, t2a, OUT
|
||||||
import translate_post as tp # carta_article_ids
|
import translate_post as tp # carta_article_ids
|
||||||
|
|
||||||
VOICE = "NicoFeadulta2026"
|
VOICE = os.environ.get("FEA_TTS_VOICE", "NicoFeadulta2026")
|
||||||
MODEL = "speech-2.8-hd"
|
MODEL = "speech-2.8-hd"
|
||||||
CONTAINER = "wordpress-web"
|
CONTAINER = "wordpress-web"
|
||||||
PROD = Path(__file__).resolve().parent.parent / "wordpress/wp-content/uploads/tts"
|
PROD = Path(__file__).resolve().parent.parent / "wordpress/wp-content/uploads/tts"
|
||||||
LOG = Path("/tmp/feadulta-tts-prod.log")
|
LOG = Path("/tmp/feadulta-tts-prod.log")
|
||||||
INTERVAL = 180 # s entre cartas exitosas (reparte el ritmo)
|
INTERVAL = 180 # s entre cartas exitosas (reparte el ritmo)
|
||||||
BACKOFF = 1800 # s de espera ante fallo de cuota antes de reintentar
|
BACKOFF = 1800 # s de espera ante errores transitorios no clasificados
|
||||||
MAX_CONSEC_FAIL = 3 # fallos seguidos → parar (cuota probablemente agotada)
|
MAX_CONSEC_FAIL = 3 # fallos seguidos no clasificados → parar
|
||||||
|
QUOTA_OR_RATE_ERRORS = {2056, 1039} # MiniMax: no reintentar en este proceso
|
||||||
MIN_CHARS = 200 # por debajo, se considera sin contenido locutable
|
MIN_CHARS = 200 # por debajo, se considera sin contenido locutable
|
||||||
|
|
||||||
# Cola de cartas a locutar. Override por entorno (FEA_TTS_CARTAS) para priorizar
|
# Cola de cartas a locutar. Override por entorno (FEA_TTS_CARTAS) para priorizar
|
||||||
@@ -52,6 +62,18 @@ def meta(pid, key):
|
|||||||
return php("getmeta", str(pid), key).stdout.strip()
|
return php("getmeta", str(pid), key).stdout.strip()
|
||||||
|
|
||||||
|
|
||||||
|
def backlog_ids(autor, desde, hasta, limite):
|
||||||
|
"""Cola del backlog de un autor, delegada a la BD (ver listpending)."""
|
||||||
|
# Voz clonada del autor, si la tiene: los locutados con otra voz también
|
||||||
|
# cuentan como pendientes. Sin clon (""), pendiente = simplemente sin audio.
|
||||||
|
voz = mm.voice_for_author(autor, "")
|
||||||
|
r = php("listpending", str(autor), str(desde), str(hasta), str(limite), voz)
|
||||||
|
if r.returncode != 0:
|
||||||
|
log(f"listpending falló (rc={r.returncode}): {r.stderr.strip()[:200]}")
|
||||||
|
return []
|
||||||
|
return [int(x) for x in r.stdout.split() if x.strip().isdigit()]
|
||||||
|
|
||||||
|
|
||||||
def build_queue():
|
def build_queue():
|
||||||
# Cola literal de IDs (ya filtrada/ordenada) para priorizar la carta nueva.
|
# Cola literal de IDs (ya filtrada/ordenada) para priorizar la carta nueva.
|
||||||
ids_override = os.environ.get("FEA_TTS_IDS", "").replace(",", " ").split()
|
ids_override = os.environ.get("FEA_TTS_IDS", "").replace(",", " ").split()
|
||||||
@@ -67,11 +89,48 @@ def build_queue():
|
|||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
|
global CARTAS
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Locuta posts ES de Fe Adulta con MiniMax; sin --ids conserva la cola programada."
|
||||||
|
)
|
||||||
|
parser.add_argument("--ids", help="CSV de IDs ES concretos, en el orden de locución deseado")
|
||||||
|
parser.add_argument("--cartas", help="CSV de cartas para construir la cola; sustituye FEA_TTS_CARTAS")
|
||||||
|
parser.add_argument("--autor", type=int,
|
||||||
|
help="WP user_id: cola del backlog de ese autor en vez de cartas")
|
||||||
|
parser.add_argument("--desde", type=int, default=0, help="año inicial del backlog (con --autor)")
|
||||||
|
parser.add_argument("--hasta", type=int, default=9999, help="año final del backlog (con --autor)")
|
||||||
|
parser.add_argument("--max", type=int, default=0,
|
||||||
|
help="para tras N audios OK en esta ejecución (0 = sin tope)")
|
||||||
|
parser.add_argument("--dry-run", action="store_true",
|
||||||
|
help="imprime la cola y sale, sin sintetizar ni gastar cuota")
|
||||||
|
parser.add_argument("--allow-default-voice", action="store_true",
|
||||||
|
help="permite Nico para autores sin voz digitalizada (desactivado por defecto)")
|
||||||
|
args = parser.parse_args()
|
||||||
|
if args.ids:
|
||||||
|
os.environ["FEA_TTS_IDS"] = args.ids
|
||||||
|
if args.cartas:
|
||||||
|
os.environ["FEA_TTS_CARTAS"] = args.cartas
|
||||||
|
CARTAS = args.cartas.replace(",", " ").split()
|
||||||
|
|
||||||
PROD.mkdir(parents=True, exist_ok=True)
|
PROD.mkdir(parents=True, exist_ok=True)
|
||||||
subprocess.run(["docker", "cp", "scripts/fea_post_io.php", f"{CONTAINER}:/tmp/fea_post_io.php"],
|
subprocess.run(["docker", "cp", "scripts/fea_post_io.php", f"{CONTAINER}:/tmp/fea_post_io.php"],
|
||||||
capture_output=True)
|
capture_output=True)
|
||||||
|
|
||||||
|
if args.autor:
|
||||||
|
# Pide holgura sobre --max: parte de la cola puede caerse por contenido corto.
|
||||||
|
limite = args.max * 3 if args.max else 0
|
||||||
|
queue = backlog_ids(args.autor, args.desde, args.hasta, limite)
|
||||||
|
origen = (f"backlog autor {args.autor} ({args.desde}-{args.hasta}), "
|
||||||
|
f"voz {mm.voice_for_author(args.autor, VOICE)}")
|
||||||
|
else:
|
||||||
queue = build_queue()
|
queue = build_queue()
|
||||||
log(f"=== INICIO orquestador TTS. Cola: {len(queue)} posts ES del gap ===")
|
origen = "cartas"
|
||||||
|
tope = f", tope {args.max} esta tanda" if args.max else ""
|
||||||
|
log(f"=== INICIO orquestador TTS. Cola: {len(queue)} posts ES [{origen}]{tope} ===")
|
||||||
|
|
||||||
|
if args.dry_run:
|
||||||
|
log("--dry-run: no sintetizo. Cola = " + (",".join(str(x) for x in queue) or "(vacía)"))
|
||||||
|
return
|
||||||
|
|
||||||
i = consec = ok = 0
|
i = consec = ok = 0
|
||||||
while i < len(queue):
|
while i < len(queue):
|
||||||
@@ -92,7 +151,15 @@ def main():
|
|||||||
i += 1
|
i += 1
|
||||||
continue
|
continue
|
||||||
|
|
||||||
voice = mm.voice_for_author(author, VOICE)
|
cloned_voice = mm.voice_for_author(author, "")
|
||||||
|
if cloned_voice:
|
||||||
|
voice = cloned_voice
|
||||||
|
elif args.allow_default_voice:
|
||||||
|
voice = VOICE
|
||||||
|
else:
|
||||||
|
log(f"#{pid}: autor sin voz digitalizada (author_id={author}); omitido")
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
rc = mm.t2a(mm.add_pauses(text), voice, MODEL, f"prod-{pid}")
|
rc = mm.t2a(mm.add_pauses(text), voice, MODEL, f"prod-{pid}")
|
||||||
if rc == 0:
|
if rc == 0:
|
||||||
src = mm.OUT / f"prod-{pid}.mp3"
|
src = mm.OUT / f"prod-{pid}.mp3"
|
||||||
@@ -104,16 +171,24 @@ def main():
|
|||||||
voice_tag = f" [{voice}]" if voice != VOICE else ""
|
voice_tag = f" [{voice}]" if voice != VOICE else ""
|
||||||
log(f"#{pid} OK «{title[:45]}»{voice_tag} → tts/{pid}.mp3 (total {ok})")
|
log(f"#{pid} OK «{title[:45]}»{voice_tag} → tts/{pid}.mp3 (total {ok})")
|
||||||
i += 1
|
i += 1
|
||||||
|
if args.max and ok >= args.max:
|
||||||
|
log(f"Tope de la tanda alcanzado ({args.max}). PARO. "
|
||||||
|
"Reanudable: la próxima ventana recalcula la cola y sigue.")
|
||||||
|
break
|
||||||
time.sleep(INTERVAL)
|
time.sleep(INTERVAL)
|
||||||
else:
|
else:
|
||||||
consec += 1
|
consec += 1
|
||||||
log(f"#{pid} FALLO rc={rc} (fallo seguido {consec}/{MAX_CONSEC_FAIL})")
|
log(f"#{pid} FALLO rc={rc} (fallo seguido {consec}/{MAX_CONSEC_FAIL})")
|
||||||
php("setflag", str(pid), "fea_audio_error", str(rc))
|
php("setflag", str(pid), "fea_audio_error", str(rc))
|
||||||
if consec >= MAX_CONSEC_FAIL:
|
if rc in QUOTA_OR_RATE_ERRORS:
|
||||||
log("Demasiados fallos seguidos → cuota agotada probablemente. PARO. "
|
log(f"MiniMax rc={rc}: cuota/rate limit explícito. PARO sin reintentar. "
|
||||||
"Reanudable: relanzar el script más tarde (salta lo ya hecho).")
|
"Reanudable: relanzar el script más tarde (salta lo ya hecho).")
|
||||||
break
|
break
|
||||||
time.sleep(BACKOFF) # reintenta el mismo post tras esperar
|
if consec >= MAX_CONSEC_FAIL:
|
||||||
|
log("Demasiados fallos seguidos no clasificados. PARO. "
|
||||||
|
"Reanudable: relanzar el script más tarde (salta lo ya hecho).")
|
||||||
|
break
|
||||||
|
time.sleep(BACKOFF) # solo errores transitorios no clasificados
|
||||||
|
|
||||||
log(f"=== FIN tanda. {ok} audios generados esta ejecución. ===")
|
log(f"=== FIN tanda. {ok} audios generados esta ejecución. ===")
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user