Compare commits
10
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
55bd7a027d | ||
|
|
37ea5e2c31 | ||
|
|
78045f2c81 | ||
|
|
4653aa5e8e | ||
|
|
956aed0971 | ||
|
|
3b2e576284 | ||
|
|
de05d746cc | ||
|
|
c677206d03 | ||
|
|
d1415d5d21 | ||
|
|
5d8bcd61ea |
@@ -0,0 +1,71 @@
|
||||
# Сессия 2026-08-17 — HTTP 502 при парсинге 19 МБ PDF = OOMKill пода
|
||||
|
||||
_Стенд: ТЕСТ, `contractor.pythonk8s.dev.nubes.ru` (приставка `dev.nubes.ru` общая для dev+test).
|
||||
instanceUid: `b4523aba-b5e6-40f1-be56-bb4d2509357c`, realm `iot-naeel`, domain `contractor`._
|
||||
|
||||
## 1. Симптом (из UI, колонка «Парсинг»)
|
||||
|
||||
Файл `0144-03-2023_отчет об оценке.pdf` (19.0 МБ) в списке **дважды**:
|
||||
- первый проход: **✗ HTTP 502**;
|
||||
- повторный: **✓ 4083 эл. (0.0с)**.
|
||||
|
||||
Остальные файлы (623 КБ, 473 КБ, 596 КБ) парсились нормально.
|
||||
Вывод: загрузка файла проходит, падает **парсинг** (тяжёлая операция в запросе).
|
||||
|
||||
## 2. Диагностика kubectl (с ВМ remote-dev = 5.172.178.213)
|
||||
|
||||
```
|
||||
kubectl get pods -n b4523aba-...
|
||||
pythonk8s-589db8db9c-qjjnc 1/1 Running 1 (5m18s ago) 17m
|
||||
```
|
||||
|
||||
`kubectl describe pod`:
|
||||
```
|
||||
Last State: Terminated
|
||||
Reason: OOMKilled
|
||||
Exit Code: 137
|
||||
Restart Count: 1
|
||||
Limits: cpu: 1, memory: 1Gi
|
||||
Requests: cpu: 1, memory: 1Gi
|
||||
```
|
||||
|
||||
`kubectl top pod` (текущий под, простое): **727Mi из 1Gi (~71%)**.
|
||||
|
||||
События:
|
||||
```
|
||||
17m Killing pod/pythonk8s-5895c478b5-7dcqq — Stopping container app
|
||||
```
|
||||
(убитый OOM-ом под; текущий `pythonk8s-589db8db9c-qjjnc` — новый).
|
||||
|
||||
Логи прежнего контейнера (`--previous`): обычные запросы (health, batch-progress,
|
||||
apply-groups, process-v2, cleanup), обрыв на 14:13:14 → OOM.
|
||||
|
||||
## 3. Вывод (по фактам)
|
||||
|
||||
**HTTP 502 = OOMKill пода (exit 137).** Синхронный парсинг 19 МБ PDF
|
||||
(`site/routes/upload_bp.py`: `parse_file` → `set_parsed` в том же запросе,
|
||||
pdfplumber → 4083 элемента) превышает лимит памяти **1Gi** → kubelet убивает под →
|
||||
шлюз не получает ответ → 502. После рестарта повторный парсинг проходит.
|
||||
|
||||
- **Это НЕ** проблема сети/размера тела (как 64KB на `services`), а **нехватка памяти**.
|
||||
- Базовое потребление уже ~71% лимита, парсинг большого PDF добивает под.
|
||||
|
||||
## 4. Связь с прежними находками
|
||||
|
||||
| Симптом | Причина | Платформа |
|
||||
|---|---|---|
|
||||
| «64KB / HTTP 000, Сеть» (исторически) | ddos-guard + таймауты Ingress + медленная обработка | `pythonk8s.services.ngcloud.ru` |
|
||||
| HTTP 502 на 19 МБ PDF (сейчас) | OOMKill: лимит 1Gi, синхронный парсинг | `pythonk8s.dev.nubes.ru` (ТЕСТ) |
|
||||
|
||||
## 5. Варианты решения (НЕ ВНЕСЕНЫ, ждут команды)
|
||||
|
||||
1. Поднять память инстанса: `clusterConfiguration.memory` 1024 → 2048 (быстрый обходной путь).
|
||||
2. Убрать синхронный парсинг из `/upload` — парсинг в фон (thread / отдельный endpoint),
|
||||
ответ сразу; снижает пик памяти в запросе (правильный фикс).
|
||||
3. Оба варианта.
|
||||
|
||||
## 6. Статус (следующий шаг)
|
||||
|
||||
Пользователь планирует **редеплой на обычном облачном кластере со стандартными настройками**
|
||||
(не на своём `iot-naeel`) — проверить, как ведёт себя парсинг 19 МБ при стандартных лимитах.
|
||||
Результат сравнить с данным диагностикой.
|
||||
@@ -92,6 +92,22 @@ server {
|
||||
client_max_body_size 100m;
|
||||
}
|
||||
|
||||
# ── VM-буфер загрузки contracts (паттерн drhider) ──────────────
|
||||
# Браузер кладёт файл сюда PUT (мимо шлюза кластера ~64КБ), бэк тянет сам.
|
||||
# Файлы: /var/www/contracts-upload/ (нужно создать, chown www-data).
|
||||
# CORS: origin фронтенда = contractor.pythonk8s.dev.nubes.ru (ingress, подтверждено).
|
||||
location /contracts-upload/ {
|
||||
alias /var/www/contracts-upload/;
|
||||
dav_methods PUT DELETE;
|
||||
create_full_put_path on;
|
||||
client_max_body_size 1024m;
|
||||
add_header Access-Control-Allow-Origin https://contractor.pythonk8s.dev.nubes.ru always;
|
||||
add_header Access-Control-Allow-Methods 'PUT, GET, OPTIONS, DELETE' always;
|
||||
add_header Access-Control-Allow-Headers 'Content-Type' always;
|
||||
add_header Access-Control-Max-Age 86400 always;
|
||||
if ($request_method = OPTIONS) { return 204; }
|
||||
}
|
||||
|
||||
listen 443 ssl; # managed by Certbot
|
||||
ssl_certificate /etc/letsencrypt/live/contracts.kube5s.ru/fullchain.pem; # managed by Certbot
|
||||
ssl_certificate_key /etc/letsencrypt/live/contracts.kube5s.ru/privkey.pem; # managed by Certbot
|
||||
|
||||
+14
-1
@@ -1,7 +1,7 @@
|
||||
"""Конфигурация приложения — все настройки в одном месте."""
|
||||
import os
|
||||
|
||||
VERSION = "2.0.5"
|
||||
VERSION = "2.0.11"
|
||||
|
||||
LLM_URL = os.getenv("LLM_API_URL", "https://api.aillm.ru/v1/chat/completions")
|
||||
LLM_KEY = os.getenv("LLM_API_KEY", "")
|
||||
@@ -9,3 +9,16 @@ LLM_MODEL = os.getenv("LLM_MODEL", "gpt-oss-120b")
|
||||
|
||||
MAX_CONTENT_LENGTH = 200 * 1024 * 1024 # 200 MB
|
||||
API_KEY = os.getenv("API_KEY", "")
|
||||
CONVERT_SERVICE_URL = os.getenv("CONVERT_SERVICE_URL", "http://containerk8s.df36c8af-1a95-4623-b551-0d37b731ccca.svc.cluster.local:5000")
|
||||
|
||||
# ── VM-буфер загрузки (паттерн drhider) ──────────────────────────────
|
||||
# Браузер кладёт файл на ВМ через WebDAV (мимо шлюза кластера ~64КБ),
|
||||
# бэк сам тянет его исходящим GET (egress без лимита).
|
||||
VM_UPLOAD_URL = os.getenv("VM_UPLOAD_URL", "https://contracts.kube5s.ru/contracts-upload/")
|
||||
# SSRF-защита: тянуть можно ТОЛЬКО с этого префикса.
|
||||
VM_UPLOAD_PREFIX = os.getenv("VM_UPLOAD_PREFIX", "https://contracts.kube5s.ru/contracts-upload/")
|
||||
# Лимит на один файл (совпадает с фронтом).
|
||||
VM_UPLOAD_MAX_BYTES = int(os.getenv("VM_UPLOAD_MAX_BYTES", str(50 * 1024 * 1024)))
|
||||
# Ретраи pull с ВМ (разовые DNS/сетевые сбои не роняют загрузку).
|
||||
PULL_RETRIES = int(os.getenv("PULL_RETRIES", "3"))
|
||||
PULL_RETRY_DELAY = 2.0
|
||||
|
||||
+227
-95
@@ -1,9 +1,10 @@
|
||||
"""Upload blueprint — загрузка, конвертация, распаковка."""
|
||||
import io, os, base64, hashlib, zipfile, tempfile, subprocess
|
||||
import io, os, time, base64, hashlib, zipfile
|
||||
import httpx
|
||||
from flask import Blueprint, request, jsonify, send_file
|
||||
from services.parse import parse_file
|
||||
from db import documents
|
||||
from config import MAX_CONTENT_LENGTH
|
||||
import config
|
||||
|
||||
upload_bp = Blueprint("upload", __name__)
|
||||
|
||||
@@ -17,9 +18,148 @@ def _check_ext(filename: str) -> str | None:
|
||||
return None
|
||||
|
||||
|
||||
def _safe_name(name: str) -> str:
|
||||
"""Санитизация имени файла: только basename, защита от path traversal."""
|
||||
name = (name or "").replace("\\", "/").rsplit("/", 1)[-1].strip()
|
||||
if not name or name in (".", ".."):
|
||||
return "file.bin"
|
||||
return name[:255]
|
||||
|
||||
|
||||
def _pull_with_retries(url: str, timeout: int = 120) -> bytes:
|
||||
"""Скачать файл с ВМ-буфера с ретраями (egress, без лимита шлюза)."""
|
||||
last: Exception | None = None
|
||||
for attempt in range(1, config.PULL_RETRIES + 1):
|
||||
try:
|
||||
resp = httpx.get(url, timeout=timeout, follow_redirects=False)
|
||||
if resp.status_code == 200:
|
||||
return resp.content
|
||||
last = Exception(f"HTTP {resp.status_code}")
|
||||
except Exception as e:
|
||||
last = e
|
||||
if attempt < config.PULL_RETRIES:
|
||||
time.sleep(config.PULL_RETRY_DELAY)
|
||||
raise last or Exception("pull failed")
|
||||
|
||||
|
||||
def _delete_from_vm(url: str) -> None:
|
||||
"""Best-effort удаление файла с ВМ-буфера."""
|
||||
try:
|
||||
httpx.delete(url, timeout=30)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
def _pull_from_ref(ref):
|
||||
"""SSRF-проверка → лимит → pull с ретраями → DELETE с ВМ. Возвращает (name, content)."""
|
||||
name = _safe_name(str(ref.get("name", "") or ""))
|
||||
size = int(ref.get("size") or 0)
|
||||
url = str(ref.get("url", "") or "")
|
||||
|
||||
if not url.startswith(config.VM_UPLOAD_PREFIX):
|
||||
raise Exception("invalid url (SSRF guard)")
|
||||
if size > config.VM_UPLOAD_MAX_BYTES:
|
||||
_delete_from_vm(url)
|
||||
raise Exception(f"file too large: {size} bytes (max {config.VM_UPLOAD_MAX_BYTES})")
|
||||
|
||||
content = _pull_with_retries(url)
|
||||
if len(content) > config.VM_UPLOAD_MAX_BYTES:
|
||||
_delete_from_vm(url)
|
||||
raise Exception("file too large after pull")
|
||||
|
||||
_delete_from_vm(url)
|
||||
return name, content
|
||||
|
||||
|
||||
def _unzip(data: bytes):
|
||||
"""Распаковать ZIP → (ok, files, error)."""
|
||||
MAX_FILES = 500
|
||||
MAX_UNCOMPRESSED = 500 * 1024 * 1024 # 500 MB
|
||||
files = []
|
||||
total = 0
|
||||
try:
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as zf:
|
||||
if len(zf.namelist()) > MAX_FILES:
|
||||
return False, None, f"too many files in ZIP (max {MAX_FILES})"
|
||||
for info in zf.infolist():
|
||||
if info.is_dir():
|
||||
continue
|
||||
name = os.path.basename(info.filename)
|
||||
if not name or ".." in name or "/" in name or "\\" in name:
|
||||
continue
|
||||
raw = zf.read(info)
|
||||
total += len(raw)
|
||||
if total > MAX_UNCOMPRESSED:
|
||||
return False, None, "total uncompressed size exceeds 500 MB"
|
||||
ext = name.rsplit(".", 1)[-1].lower() if "." in name else ""
|
||||
files.append({
|
||||
"filename": name,
|
||||
"ext": ext,
|
||||
"size": len(raw),
|
||||
"data_b64": base64.b64encode(raw).decode(),
|
||||
})
|
||||
except zipfile.BadZipFile:
|
||||
return False, None, "invalid ZIP archive"
|
||||
return True, files, None
|
||||
|
||||
|
||||
def _convert(filename: str, data: bytes) -> bytes:
|
||||
""".doc → .docx через внешний libreoffice-сервис. Возвращает docx-байты."""
|
||||
try:
|
||||
resp = httpx.post(
|
||||
config.CONVERT_SERVICE_URL + "/convert",
|
||||
files={"file": (filename, data, "application/msword")},
|
||||
timeout=120,
|
||||
)
|
||||
except httpx.TimeoutException:
|
||||
raise Exception("conversion timeout")
|
||||
if resp.status_code != 200:
|
||||
try:
|
||||
err = resp.json().get("error", "conversion failed")
|
||||
except Exception:
|
||||
err = "conversion failed"
|
||||
raise Exception(err)
|
||||
return resp.content
|
||||
|
||||
|
||||
def _store_and_parse(filename: str, data: bytes, batch_id, contract_id, zip_source=None, mime_type="application/octet-stream"):
|
||||
"""Общая логика: дедуп → insert в БД → авто-парсинг. Возвращает dict-результат."""
|
||||
content_hash = hashlib.sha256(data).hexdigest()[:16]
|
||||
|
||||
# Дедупликация по хешу
|
||||
if batch_id:
|
||||
existing = documents.get_by_hash(batch_id, content_hash)
|
||||
if existing:
|
||||
return {"ok": False, "error": "duplicate", "doc_id": existing["id"], "duplicate_of": True}
|
||||
|
||||
doc = documents.insert(
|
||||
filename=filename,
|
||||
mime_type=mime_type,
|
||||
original_bytes=base64.b64encode(data).decode(),
|
||||
batch_id=batch_id,
|
||||
zip_source=zip_source,
|
||||
content_hash=content_hash,
|
||||
)
|
||||
|
||||
# Авто-парсинг
|
||||
try:
|
||||
result = parse_file(filename, data)
|
||||
if result["status"] == "parsed":
|
||||
documents.set_parsed(doc["id"], result["elements"])
|
||||
parsed = {"status": "parsed", "element_count": result.get("element_count", 0)}
|
||||
else:
|
||||
documents.set_error(doc["id"], result.get("error", "parse failed"))
|
||||
parsed = {"status": "error", "error": result.get("error", "parse failed")}
|
||||
except Exception as e:
|
||||
documents.set_error(doc["id"], str(e))
|
||||
parsed = {"status": "error", "error": str(e)}
|
||||
|
||||
return {"ok": True, "doc_id": doc["id"], "contract_id": contract_id, "parsed": parsed}
|
||||
|
||||
|
||||
@upload_bp.route("/upload", methods=["POST"])
|
||||
def upload():
|
||||
"""Загрузка одного файла + авто-парсинг → БД."""
|
||||
"""Загрузка одного файла + авто-парсинг → БД (прямой multipart)."""
|
||||
f = request.files.get("files")
|
||||
if not f:
|
||||
return jsonify(ok=False, error="no file"), 400
|
||||
@@ -29,117 +169,109 @@ def upload():
|
||||
return jsonify(ok=False, error=err), 400
|
||||
|
||||
data = f.read()
|
||||
content_hash = hashlib.sha256(data).hexdigest()[:16]
|
||||
batch_id = request.form.get("batch_id")
|
||||
zip_source = request.form.get("zip_source")
|
||||
|
||||
# Дедупликация по хешу
|
||||
if batch_id:
|
||||
existing = documents.get_by_hash(batch_id, content_hash)
|
||||
if existing:
|
||||
return jsonify(ok=False, error="duplicate", doc_id=existing["id"])
|
||||
|
||||
doc = documents.insert(
|
||||
filename=f.filename,
|
||||
mime_type=f.content_type or "application/octet-stream",
|
||||
original_bytes=base64.b64encode(data).decode(),
|
||||
batch_id=batch_id,
|
||||
zip_source=zip_source,
|
||||
content_hash=content_hash,
|
||||
)
|
||||
|
||||
# Авто-парсинг
|
||||
try:
|
||||
result = parse_file(f.filename, data)
|
||||
if result["status"] == "parsed":
|
||||
documents.set_parsed(doc["id"], result["elements"])
|
||||
else:
|
||||
documents.set_error(doc["id"], result.get("error", "parse failed"))
|
||||
except Exception as e:
|
||||
documents.set_error(doc["id"], str(e))
|
||||
result = {"status": "error", "error": str(e)}
|
||||
|
||||
contract_id = request.form.get("contract_id")
|
||||
return jsonify(
|
||||
ok=True,
|
||||
doc_id=doc["id"],
|
||||
contract_id=contract_id,
|
||||
parsed={"status": result["status"], "element_count": result.get("element_count", 0)},
|
||||
)
|
||||
|
||||
result = _store_and_parse(f.filename, data, batch_id, contract_id, zip_source, f.content_type or "application/octet-stream")
|
||||
if not result["ok"]:
|
||||
return jsonify(ok=result["ok"], error=result.get("error"), doc_id=result.get("doc_id")), 200
|
||||
return jsonify(ok=True, doc_id=result["doc_id"], contract_id=result["contract_id"], parsed=result["parsed"])
|
||||
|
||||
|
||||
@upload_bp.route("/api/upload_refs", methods=["POST"])
|
||||
def upload_refs():
|
||||
"""Загрузка через ВМ-буфер (паттерн drhider): бэк тянет файлы с ВМ.
|
||||
|
||||
Тело (маленькое, <64КБ): {batch_id, contract_id, zip_source, files:[{name,size,url}]}.
|
||||
Для каждой ссылки: SSRF-проверка → лимит → pull с ретраями → store+parse → DELETE с ВМ.
|
||||
"""
|
||||
data = request.get_json(silent=True) or {}
|
||||
files = data.get("files") or []
|
||||
batch_id = data.get("batch_id")
|
||||
contract_id = data.get("contract_id")
|
||||
zip_source = data.get("zip_source")
|
||||
|
||||
if not files:
|
||||
return jsonify(ok=False, error="no files"), 400
|
||||
|
||||
results = []
|
||||
for ref in files:
|
||||
ref_name = _safe_name(str(ref.get("name", "") or ""))
|
||||
try:
|
||||
name, content = _pull_from_ref(ref)
|
||||
except Exception as e:
|
||||
results.append({"name": ref_name, "ok": False, "error": str(e)})
|
||||
continue
|
||||
stored = _store_and_parse(name, content, batch_id, contract_id, zip_source)
|
||||
results.append({"name": name, **stored})
|
||||
|
||||
return jsonify(ok=True, results=results)
|
||||
|
||||
|
||||
@upload_bp.route("/convert-doc", methods=["POST"])
|
||||
def convert_doc():
|
||||
""".doc → .docx через libreoffice (без сохранения на диск)."""
|
||||
""".doc → .docx через внешний libreoffice-сервис (прямой multipart)."""
|
||||
f = request.files.get("files")
|
||||
if not f:
|
||||
return jsonify(ok=False, error="no file"), 400
|
||||
|
||||
data = f.read()
|
||||
doc_path = None
|
||||
tmpdir = None
|
||||
try:
|
||||
with tempfile.NamedTemporaryFile(suffix=".doc", delete=False) as tmp:
|
||||
tmp.write(data)
|
||||
doc_path = tmp.name
|
||||
tmpdir = tempfile.mkdtemp()
|
||||
subprocess.run(
|
||||
["libreoffice", "--headless", "--convert-to", "docx", "--outdir", tmpdir, doc_path],
|
||||
timeout=30, capture_output=True,
|
||||
)
|
||||
docx_files = [x for x in os.listdir(tmpdir) if x.endswith(".docx")]
|
||||
if docx_files:
|
||||
with open(os.path.join(tmpdir, docx_files[0]), "rb") as out:
|
||||
content = _convert(f.filename, f.read())
|
||||
except Exception as e:
|
||||
return jsonify(ok=False, error=str(e)), 500
|
||||
return send_file(
|
||||
io.BytesIO(out.read()),
|
||||
io.BytesIO(content),
|
||||
mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document",
|
||||
)
|
||||
|
||||
|
||||
@upload_bp.route("/api/convert_refs", methods=["POST"])
|
||||
def convert_refs():
|
||||
""".doc → .docx через ВМ-буфер (паттерн drhider): pull .doc с ВМ → конвертация → docx."""
|
||||
data = request.get_json(silent=True) or {}
|
||||
files = data.get("files") or []
|
||||
if not files:
|
||||
return jsonify(ok=False, error="no files"), 400
|
||||
ref = files[0]
|
||||
try:
|
||||
name, content = _pull_from_ref(ref)
|
||||
except Exception as e:
|
||||
return jsonify(ok=False, error=str(e)), 400
|
||||
try:
|
||||
docx = _convert(name, content)
|
||||
except Exception as e:
|
||||
return jsonify(ok=False, error=str(e)), 500
|
||||
return send_file(
|
||||
io.BytesIO(docx),
|
||||
mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document",
|
||||
)
|
||||
return jsonify(ok=False, error="conversion produced no output"), 500
|
||||
finally:
|
||||
if doc_path and os.path.exists(doc_path):
|
||||
os.unlink(doc_path)
|
||||
if tmpdir and os.path.exists(tmpdir):
|
||||
for x in os.listdir(tmpdir):
|
||||
os.unlink(os.path.join(tmpdir, x))
|
||||
os.rmdir(tmpdir)
|
||||
|
||||
|
||||
@upload_bp.route("/unzip-upload", methods=["POST"])
|
||||
def unzip_upload():
|
||||
"""Распаковать ZIP → список файлов (base64 для фронтенда)."""
|
||||
"""Распаковать ZIP → список файлов (base64 для фронтенда, прямой multipart)."""
|
||||
f = request.files.get("files")
|
||||
if not f:
|
||||
return jsonify(ok=False, error="no file"), 400
|
||||
|
||||
data = f.read()
|
||||
MAX_FILES = 500
|
||||
MAX_UNCOMPRESSED = 500 * 1024 * 1024 # 500 MB
|
||||
|
||||
files = []
|
||||
total = 0
|
||||
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as zf:
|
||||
if len(zf.namelist()) > MAX_FILES:
|
||||
return jsonify(ok=False, error=f"too many files in ZIP (max {MAX_FILES})"), 400
|
||||
|
||||
for info in zf.infolist():
|
||||
if info.is_dir():
|
||||
continue
|
||||
name = os.path.basename(info.filename)
|
||||
if not name or ".." in name or "/" in name or "\\" in name:
|
||||
continue
|
||||
|
||||
raw = zf.read(info)
|
||||
total += len(raw)
|
||||
if total > MAX_UNCOMPRESSED:
|
||||
return jsonify(ok=False, error="total uncompressed size exceeds 500 MB"), 400
|
||||
|
||||
ext = name.rsplit(".", 1)[-1].lower() if "." in name else ""
|
||||
files.append({
|
||||
"filename": name,
|
||||
"ext": ext,
|
||||
"size": len(raw),
|
||||
"data_b64": base64.b64encode(raw).decode(),
|
||||
})
|
||||
|
||||
ok, files, err = _unzip(f.read())
|
||||
if not ok:
|
||||
return jsonify(ok=False, error=err), 400
|
||||
return jsonify(ok=True, files=files)
|
||||
|
||||
|
||||
@upload_bp.route("/api/unzip_refs", methods=["POST"])
|
||||
def unzip_refs():
|
||||
"""Распаковать ZIP через ВМ-буфер (паттерн drhider): pull ZIP с ВМ → распаковка."""
|
||||
data = request.get_json(silent=True) or {}
|
||||
files = data.get("files") or []
|
||||
if not files:
|
||||
return jsonify(ok=False, error="no files"), 400
|
||||
ref = files[0]
|
||||
try:
|
||||
name, content = _pull_from_ref(ref)
|
||||
except Exception as e:
|
||||
return jsonify(ok=False, error=str(e)), 400
|
||||
ok, unzipped, err = _unzip(content)
|
||||
if not ok:
|
||||
return jsonify(ok=False, error=err), 400
|
||||
return jsonify(ok=True, files=unzipped)
|
||||
|
||||
@@ -4,6 +4,9 @@ var VM_API = '';
|
||||
var UPLOAD_URL = '/upload';
|
||||
var CONVERT_URL = '/convert-doc';
|
||||
var UNZIP_URL = '/unzip-upload';
|
||||
// ВМ-буфер загрузки (паттерн drhider): браузер кладёт файл сюда (WebDAV, мимо шлюза),
|
||||
// бэк сам тянет его по /api/upload_refs. Origin должен быть в CORS на nginx ВМ.
|
||||
var VM_UPLOAD_URL = 'https://contracts.kube5s.ru/contracts-upload/';
|
||||
// SITE_URL удалён (Фаза 4) — не использовался
|
||||
|
||||
var fileInput = document.getElementById('fileInput');
|
||||
|
||||
+55
-4
@@ -1,4 +1,55 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 32 32">
|
||||
<rect width="32" height="32" rx="4" fill="#001C34"/>
|
||||
<text x="16" y="24" text-anchor="middle" font-size="22" font-weight="bold" fill="white" font-family="sans-serif">N</text>
|
||||
</svg>
|
||||
<?xml version="1.0" encoding="UTF-8" standalone="no"?>
|
||||
<svg
|
||||
width="57"
|
||||
height="57"
|
||||
xml:space="preserve"
|
||||
overflow="hidden"
|
||||
version="1.1"
|
||||
id="svg8"
|
||||
sodipodi:docname="U_v3.svg"
|
||||
inkscape:version="1.3.2 (091e20e, 2023-11-25, custom)"
|
||||
xmlns:inkscape="http://www.inkscape.org/namespaces/inkscape"
|
||||
xmlns:sodipodi="http://sodipodi.sourceforge.net/DTD/sodipodi-0.dtd"
|
||||
xmlns="http://www.w3.org/2000/svg"
|
||||
xmlns:svg="http://www.w3.org/2000/svg"><sodipodi:namedview
|
||||
id="namedview8"
|
||||
pagecolor="#ffffff"
|
||||
bordercolor="#000000"
|
||||
borderopacity="0.25"
|
||||
inkscape:showpageshadow="2"
|
||||
inkscape:pageopacity="0.0"
|
||||
inkscape:pagecheckerboard="0"
|
||||
inkscape:deskcolor="#d1d1d1"
|
||||
inkscape:zoom="16.226667"
|
||||
inkscape:cx="27.146672"
|
||||
inkscape:cy="28.317584"
|
||||
inkscape:window-width="2560"
|
||||
inkscape:window-height="1494"
|
||||
inkscape:window-x="-11"
|
||||
inkscape:window-y="-11"
|
||||
inkscape:window-maximized="1"
|
||||
inkscape:current-layer="svg8" /><defs
|
||||
id="defs2"><clipPath
|
||||
id="clip0"><rect
|
||||
x="652"
|
||||
y="420"
|
||||
width="64"
|
||||
height="75"
|
||||
id="rect1" /></clipPath><clipPath
|
||||
id="clip1"><rect
|
||||
x="652"
|
||||
y="420"
|
||||
width="64"
|
||||
height="70"
|
||||
id="rect2" /></clipPath></defs><g
|
||||
clip-path="url(#clip0)"
|
||||
transform="matrix(1.0135748,0,0,1.0135748,-664.97034,-440.30567)"
|
||||
id="g8"><g
|
||||
clip-path="url(#clip1)"
|
||||
id="g7"><g
|
||||
id="g6"><path
|
||||
d="m 114.028,43.1492 v 7.6592 c 0,3.7843 -3.08,6.8632 -6.866,6.8632 L 85.3367,57.5419 c -3.7853,0 -6.8655,-3.0805 -6.8655,-6.8643 v -2.5279 l 0.0555,0.009 V 15.8148 l -10.3823,2.0127 0.0045,3.8772 -0.0045,0.0015 v 28.971 c 0,9.481 7.7132,17.1923 17.1867,17.1923 l 21.8359,0.1313 c 9.474,0 17.187,-7.7117 17.187,-17.1928 v -2.6541 l 0.027,0.0045 V 15.8145 l -10.382,2.0126 0.028,25.3214 z"
|
||||
fill="#001c34"
|
||||
fill-rule="evenodd"
|
||||
transform="matrix(1,0,0,1.01337,587.92,420.059)"
|
||||
id="path6" /></g></g></g></svg>
|
||||
|
||||
|
Before Width: | Height: | Size: 246 B After Width: | Height: | Size: 2.0 KiB |
+103
-27
@@ -12,6 +12,8 @@
|
||||
* - statusToHTML() — рендер статуса (↑ N%, ✓)
|
||||
* - renderFiles() — рендер всей таблицы
|
||||
|
||||
var CONVERT_URL = '/convert-doc';
|
||||
|
||||
/**
|
||||
* statusToHTML(status) — Чистая функция: структура → HTML (Фаза 1).
|
||||
*
|
||||
@@ -26,6 +28,7 @@ function statusToHTML(st) {
|
||||
if (!st || !st.kind) return '';
|
||||
switch (st.kind) {
|
||||
case 'connecting': return '⏳ соединение... ' + (st.elapsed || 0) + 'с';
|
||||
case 'converting': return '⏳ конвертация... ' + (st.elapsed || 0) + 'с';
|
||||
case 'uploading': return '⏳ отправка... ' + (st.elapsed || 0) + 'с';
|
||||
case 'uploaded': return '<span class="status-ok">✓</span>';
|
||||
case 'unzipping': return '⏳ распаковка...';
|
||||
@@ -51,7 +54,7 @@ function statusToHTML(st) {
|
||||
*/
|
||||
function renderFiles(state) {
|
||||
if (state.files.length === 0) {
|
||||
fileTable.innerHTML = '<tr class="empty-row"><td colspan="7">Нет файлов — выберите .docx / .pdf</td></tr>';
|
||||
fileTable.innerHTML = '<tr class="empty-row"><td colspan="7">Нет файлов — выберите .doc / .docx / .pdf</td></tr>';
|
||||
lucide.createIcons();
|
||||
return;
|
||||
}
|
||||
@@ -204,21 +207,61 @@ window.toggleClassifyDetail = async function(i) {
|
||||
};
|
||||
|
||||
/**
|
||||
* ⛔ НЕ МЕНЯТЬ ⛔ uploadFile — fetch-загрузка с честным счётчиком времени.
|
||||
* convertDoc(file, onProgress) — .doc → .docx через ВМ-буфер (паттерн drhider).
|
||||
* Фаза 1: PUT .doc на ВМ. Фаза 2: /api/convert_refs → pull → конвертация → .docx.
|
||||
*/
|
||||
function convertDoc(file, onProgress) {
|
||||
var startTime = Date.now();
|
||||
if (onProgress) onProgress({ kind: 'converting', elapsed: 0 });
|
||||
var token = crypto.randomUUID();
|
||||
var vmUrl = VM_UPLOAD_URL + token + '_0';
|
||||
|
||||
var timer = setInterval(function() {
|
||||
var elapsed = Math.floor((Date.now() - startTime) / 1000);
|
||||
if (onProgress) onProgress({ kind: 'converting', elapsed: elapsed });
|
||||
}, 1000);
|
||||
|
||||
return fetch(vmUrl, { method: 'PUT', body: file, headers: { 'Content-Type': 'application/octet-stream' } })
|
||||
.then(function(r) {
|
||||
if (!r.ok) throw new Error('Конвертация: VM upload HTTP ' + r.status);
|
||||
return fetch('/api/convert_refs?_=' + Date.now(), {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ files: [{ name: file.name, size: file.size, url: vmUrl }] })
|
||||
});
|
||||
})
|
||||
.then(function(r) {
|
||||
if (!r.ok) throw new Error('Конвертация: HTTP ' + r.status);
|
||||
return r.blob();
|
||||
})
|
||||
.then(function(blob) {
|
||||
clearInterval(timer);
|
||||
if (blob.size === 0) throw new Error('Конвертация: пустой результат');
|
||||
return new File([blob], file.name.replace(/\.doc$/i, '.docx'), {
|
||||
type: 'application/vnd.openxmlformats-officedocument.wordprocessingml.document',
|
||||
lastModified: Date.now()
|
||||
});
|
||||
})
|
||||
.catch(function(e) {
|
||||
clearInterval(timer);
|
||||
if (e.message === 'Failed to fetch' || e.name === 'TypeError') throw new Error('Конвертация: Сеть');
|
||||
throw e;
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* ⛔ НЕ МЕНЯТЬ БЕЗ РАЗРЕШЕНИЯ НАЕЛЯ ⛔ uploadFile — загрузка через ВМ-буфер (паттерн drhider).
|
||||
*
|
||||
* v2.0.3: вместо фейкового ↑N% — честный счётчик ⏳ соединение... Nс → ⏳ отправка... Nс.
|
||||
* fetch() не даёт реальный upload progress, поэтому считаем секунды.
|
||||
* Кэш-бастинг: ?_=Date.now()
|
||||
* Фаза 1: PUT файла на ВМ (WebDAV /contracts-upload/, мимо шлюза кластера ~64КБ).
|
||||
* Фаза 2: POST /api/upload_refs {files:[{name,size,url}]} — бэк сам тянет файл с ВМ.
|
||||
* Честный счётчик времени: ⏳ соединение... Nс → ⏳ отправка... Nс (fetch не даёт progress).
|
||||
*/
|
||||
function uploadFile(file, onProgress, zipSource) {
|
||||
var startTime = Date.now();
|
||||
var phase = 'connecting'; // connecting → uploading
|
||||
if (onProgress) onProgress({ kind: 'connecting', elapsed: 0 });
|
||||
var fd = new FormData();
|
||||
fd.append('files', file, file.name);
|
||||
if (state.contractId) fd.append('contract_id', state.contractId);
|
||||
fd.append('batch_id', state.batchId);
|
||||
if (zipSource) fd.append('zip_source', zipSource);
|
||||
var token = crypto.randomUUID();
|
||||
var vmUrl = VM_UPLOAD_URL + token + '_0';
|
||||
|
||||
// Честный счётчик: каждую секунду обновляем elapsed
|
||||
var timer = setInterval(function() {
|
||||
@@ -226,17 +269,33 @@ function uploadFile(file, onProgress, zipSource) {
|
||||
if (onProgress) onProgress({ kind: phase, elapsed: elapsed });
|
||||
}, 1000);
|
||||
|
||||
return fetch(UPLOAD_URL + '?_=' + Date.now(), { method: 'POST', body: fd })
|
||||
// Фаза 1: PUT файла на ВМ-буфер
|
||||
return fetch(vmUrl, { method: 'PUT', body: file, headers: { 'Content-Type': 'application/octet-stream' } })
|
||||
.then(function(r) {
|
||||
if (!r.ok) throw new Error('VM upload HTTP ' + r.status);
|
||||
phase = 'uploading';
|
||||
// Фаза 2: refs на бэкенд → pull с ВМ
|
||||
var body = { batch_id: state.batchId, files: [{ name: file.name, size: file.size, url: vmUrl }] };
|
||||
if (state.contractId) body.contract_id = state.contractId;
|
||||
if (zipSource) body.zip_source = zipSource;
|
||||
return fetch('/api/upload_refs?_=' + Date.now(), {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify(body)
|
||||
});
|
||||
})
|
||||
.then(function(r) {
|
||||
if (!r.ok) throw new Error('HTTP ' + r.status);
|
||||
return r.json();
|
||||
})
|
||||
.then(function(data) {
|
||||
clearInterval(timer);
|
||||
if (data.ok) {
|
||||
if (data.ok && data.results && data.results.length === 1) {
|
||||
var res = data.results[0];
|
||||
if (!res.ok) throw new Error(res.error || 'Неизвестная ошибка');
|
||||
var elapsed = Math.floor((Date.now() - startTime) / 1000);
|
||||
if (onProgress) onProgress({ kind: 'uploading', elapsed: elapsed });
|
||||
return data;
|
||||
return res; // {ok, doc_id, contract_id, parsed}
|
||||
}
|
||||
throw new Error(data.error || 'Неизвестная ошибка');
|
||||
})
|
||||
@@ -336,19 +395,18 @@ async function addZipFile(file) {
|
||||
render(state);
|
||||
|
||||
try {
|
||||
// Шаг 1: распаковать ZIP на бэкенде
|
||||
var zipResp = await new Promise(function(resolve, reject) {
|
||||
var xhr = new XMLHttpRequest();
|
||||
xhr.open('POST', UNZIP_URL);
|
||||
xhr.responseType = 'json';
|
||||
xhr.onload = function() { resolve(xhr.response); };
|
||||
xhr.onerror = function() { reject(new Error('Сеть')); };
|
||||
xhr.ontimeout = function() { reject(new Error('Таймаут')); };
|
||||
xhr.timeout = 60000;
|
||||
var fd = new FormData();
|
||||
fd.append('files', file);
|
||||
xhr.send(fd);
|
||||
// Шаг 1: распаковать ZIP через ВМ-буфер (паттерн drhider)
|
||||
var token = crypto.randomUUID();
|
||||
var vmUrl = VM_UPLOAD_URL + token + '_0';
|
||||
var putResp = await fetch(vmUrl, { method: 'PUT', body: file, headers: { 'Content-Type': 'application/octet-stream' } });
|
||||
if (!putResp.ok) throw new Error('unzip failed: VM upload HTTP ' + putResp.status);
|
||||
var refsResp = await fetch('/api/unzip_refs?_=' + Date.now(), {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ files: [{ name: file.name, size: file.size, url: vmUrl }] })
|
||||
});
|
||||
if (!refsResp.ok) throw new Error('unzip failed: HTTP ' + refsResp.status);
|
||||
var zipResp = await refsResp.json();
|
||||
if (!zipResp.ok || !zipResp.files) throw new Error('unzip failed');
|
||||
|
||||
// Подтверждение: показать первые 10 файлов + итог
|
||||
@@ -430,8 +488,17 @@ async function addRegularFile(file, zipSource) {
|
||||
render(state);
|
||||
|
||||
try {
|
||||
// .doc → конвертация через libreoffice-сервис
|
||||
var uploadTarget = file;
|
||||
if (file.name.toLowerCase().endsWith('.doc')) {
|
||||
uploadTarget = await convertDoc(file, function(st) {
|
||||
state.files[rowIdx].status = st;
|
||||
render(state);
|
||||
});
|
||||
state.files[rowIdx].name = uploadTarget.name;
|
||||
}
|
||||
// fetch-загрузка с честным счётчиком времени
|
||||
var resp = await uploadFile(file, function(st) {
|
||||
var resp = await uploadFile(uploadTarget, function(st) {
|
||||
state.files[rowIdx].status = st;
|
||||
render(state);
|
||||
}, zipSource);
|
||||
@@ -523,7 +590,16 @@ async function onFilesSelected(newFiles) {
|
||||
delete entry._pendingFile;
|
||||
var f = pendingFile;
|
||||
try {
|
||||
var resp = await uploadFile(f, function(st) {
|
||||
// .doc → конвертация через libreoffice-сервис
|
||||
var uploadTarget = f;
|
||||
if (f.name.toLowerCase().endsWith('.doc')) {
|
||||
uploadTarget = await convertDoc(f, function(st) {
|
||||
entry.status = st;
|
||||
render(state);
|
||||
});
|
||||
entry.name = uploadTarget.name;
|
||||
}
|
||||
var resp = await uploadFile(uploadTarget, function(st) {
|
||||
entry.status = st;
|
||||
render(state);
|
||||
});
|
||||
|
||||
@@ -73,7 +73,7 @@
|
||||
<body>
|
||||
<div class="topbar">
|
||||
<img src="/static/logo.svg" alt="Nubes">
|
||||
<span class="title">Сверка договоров — LLM AI-driven Event Sourcing <span style="font-weight:400;color:var(--muted);font-size:12px;">v2.0.5</span></span>
|
||||
<span class="title">Сверка договоров — LLM AI-driven Event Sourcing <span style="font-weight:400;color:var(--muted);font-size:12px;">v2.0.11</span></span>
|
||||
<div id="pipelineStepper" style="display:flex;gap:8px;font-size:11px;align-items:center;color:var(--muted);">
|
||||
<span id="stepUpload">○ Загрузка</span><span>→</span>
|
||||
<span id="stepClassify">○ Классификация</span><span>→</span>
|
||||
@@ -90,7 +90,7 @@
|
||||
Загрузка договоров/приложений/спецификаций
|
||||
</div>
|
||||
<div class="card-body">
|
||||
<input type="file" id="fileInput" accept=".docx,.pdf,.zip" multiple style="margin-bottom:6px;width:100%;">
|
||||
<input type="file" id="fileInput" accept=".doc,.docx,.pdf,.zip" multiple style="margin-bottom:6px;width:100%;">
|
||||
<div style="text-align:right;font-size:11px;color:var(--muted);margin-bottom:6px;">Порядок определяется автоматически при классификации</div>
|
||||
<div style="font-size:11px;color:var(--muted);margin-bottom:6px;">⚠ При совпадении имён — запрос на перезапись (OK / Отмена). Файлы из ZIP-архивов загружаются через тот же поток.</div>
|
||||
|
||||
@@ -216,11 +216,11 @@
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<script src="/static/state.js?v=2.0.5"></script>
|
||||
<script src="/static/app_utils.js?v=2.0.5"></script>
|
||||
<script src="/static/files.js?v=2.0.5"></script>
|
||||
<script src="/static/groups.js?v=2.0.5"></script>
|
||||
<script src="/static/compare.js?v=2.0.5"></script>
|
||||
<script src="/static/app.js?v=2.0.5"></script>
|
||||
<script src="/static/state.js?v=2.0.8"></script>
|
||||
<script src="/static/app_utils.js?v=2.0.8"></script>
|
||||
<script src="/static/files.js?v=2.0.11"></script>
|
||||
<script src="/static/groups.js?v=2.0.8"></script>
|
||||
<script src="/static/compare.js?v=2.0.8"></script>
|
||||
<script src="/static/app.js?v=2.0.10"></script>
|
||||
</body>
|
||||
</html>
|
||||
|
||||
Reference in New Issue
Block a user