Hermes as the AI gateway β "catch everything (AI)" β
Centralize all LLM generation behind Hermes so model selection, knowledge/RAG, logging, rate-limits, the shadow+eval gate, and data-residency are enforced in one place. Scope note: canonical source lives here under
services/hermes-server/andsupabase/functions/_shared/aiClient.ts. The Flutter client is gateway-agnostic (it streams the OpenAI-SSE shape). Pairs withHERMES_MODEL_CONFIG.md.
Principle β
Edge functions stay the entry points (stable contracts, auth, data fetching). Their LLM step delegates to Hermes. Hermes is the AI brain, not a proxy for all traffic.
app / cron ββΆ edge fn (auth + data fetch) ββΆ Hermes gateway ββΆ provider (CF AI Gateway ββΆ Gemini/β¦)
β model pin (hermes_config)
β RAG (hermes_embeddings)
β log (ai_usage_logs) + rate-limit (ai_usage_budgets)
β shadow/eval (vs model.*.target)
ββ (Hermes down?) edge fn falls back to direct gatewayIn scope (route through Hermes) vs not β
AI/LLM functions (~18 β route): ai-assistant, ai-regression, ai-valuation-analysis, ai-valuation-report, ai-data-analyst, ai-market-intelligence, ai-marketing-caption, ai-property-search, generate-market-report, generate-property-image, whatsapp-ai-chatbot, whatsapp-summarize, smart-whatsapp-alerts, suggest-price, suggest-parameters, classify-document, scan-document, research-rental-rates.
NOT in scope (transport / data / auth β leave as-is): mailbox-api, gmail-import-*, send-email, paci-*, import-*, all whatsapp-* webhooks/senders, process-payment, webauthn-*, push, cron reminders, etc.
The enabler β
Most AI functions already share _shared/aiClient.ts (callAI). Pointing that one helper at Hermes migrates most of them at once. Exceptions that bypass the shared helper and need individual work:
generate-market-reportβ calls Gemini directly with tool-calling.ai-data-analyst,ai-market-intelligenceβ streaming (need Hermes SSE).
Gateway contract (Flutter-owned Hermes service) β
POST /v1/completions # non-streaming, OpenAI-compatible
POST /v1/completions/stream # SSE: data: {choices:[{delta:{content}}]}
Authorization: Bearer <hermes service key> # hermes_api_keys (service-to-service)
body: {
messages: [{ role, content }],
role: "chat" | "reasoning", # -> hermes_config model.chat / model.reasoning
response_format?: "json_object" | "text",
tools?, tool_choice?, # passthrough for generate-market-report
temperature?, max_tokens?,
context_hint?: string, # opt-in RAG selector (use_case / entity)
metadata: { function_name, use_case?, user_id? } # for logging + rate-limit
}
returns: { choices:[{ message:{ content } }], model, usage:{ prompt_tokens, completion_tokens, total_tokens } }Per call, Hermes:
- resolves
model = hermes_config["model." + role](the central pin); - if a
model.<role>.targetshadow is enabled β also calls it, writes a side-by-side row (reuse thewhatsapp_shadow_repliespattern); - optionally injects RAG context from
hermes_embeddings(bycontext_hint/use_case); - enforces rate-limit +
ai_usage_budgetskeyed onmetadata.user_id; - calls the provider via the Cloudflare AI Gateway with
model; - logs
ai_usage_logs{ model, function_name, user_id, prompt/completion/ total tokens, estimated_cost_usd, duration_ms, request_id, use_case }; - returns the OpenAI-compatible shape.
The callAI β Hermes shim (_shared/aiClient.ts) β
import { getModel } from "./modelConfig.ts"; // hermes_config reader
const HERMES = Deno.env.get("HERMES_URL") ?? "https://hermes.aldilaijan.com/api";
const HERMES_KEY = Deno.env.get("HERMES_SERVICE_KEY")!;
const GATEWAY_ON = (Deno.env.get("AI_GATEWAY_ENABLED") ?? "true") === "true";
export async function callAI(o: {
messages: { role: string; content: string }[];
role?: "chat" | "reasoning";
responseFormat?: "json_object" | "text";
temperature?: number;
maxTokens?: number;
meta: { functionName: string; useCase?: string; userId?: string };
}): Promise<{ content: string; model: string }> {
if (GATEWAY_ON) {
try {
const r = await fetch(`${HERMES}/v1/completions`, {
method: "POST",
headers: { "content-type": "application/json", authorization: `Bearer ${HERMES_KEY}` },
body: JSON.stringify({
messages: o.messages,
role: o.role ?? "chat",
response_format: o.responseFormat,
temperature: o.temperature,
max_tokens: o.maxTokens,
metadata: o.meta,
}),
signal: AbortSignal.timeout(45_000),
});
if (r.ok) {
const j = await r.json();
return { content: j.choices[0].message.content, model: j.model };
}
console.warn(`callAI: Hermes ${r.status} β falling back to direct`);
} catch (e) {
console.warn("callAI: Hermes unreachable β falling back to direct:", e);
}
}
// Fallback: existing direct provider path, using the SAME pinned model so
// behaviour matches. (Keep the current callProviderDirect implementation.)
return callProviderDirect({ ...o, model: await getModel(o.role ?? "chat") });
}A streamAI twin does the same against /v1/completions/stream, falling back to the current direct-SSE path. Toggling AI_GATEWAY_ENABLED=false disables routing instantly (kill-switch) β also mirror it as a hermes_config key (ai.gateway.enabled) so it's flippable without redeploy.
Phased migration β
- Stand up
/v1/completions+/streamin Hermes: readhermes_config, call the CF AI Gateway, writeai_usage_logs. AddHERMES_SERVICE_KEY. - Pilot one function β
ai-valuation-analysisβ throughcallAIβHermes. Shadow vs direct; compare quality / cost / latency fromai_usage_logs. - Flip the shared
callAIβ Hermes (most functions migrate at once). Keep the fallback on; watchai_usage_logs.error_message+ p95 latency. - Exceptions: refactor
generate-market-report(tool-calling passthrough) and the streamingai-data-analyst/ai-market-intelligenceonto/stream. - Turn on RAG-context injection + the shadow/eval gate uniformly.
Rollback at every step: AI_GATEWAY_ENABLED=false (or the hermes_config key) reverts to direct calls in one flip; the per-call fallback already degrades gracefully on a Hermes blip.
Why it's worth it β
- One model pin (
hermes_config) actually governs all AI, not just chat. - One
ai_usage_logsstream β one cost/latency/quality dashboard. - The knowledge/RAG layer + house voice apply to every AI surface.
- One place to enforce data-residency (what PII β incl. anything touching
registry_votersβ reaches which provider).
Trade-offs (design for these) β
- SPOF β Hermes HA + the per-call fallback above (degrade, don't die).
- Latency β thin pass-through for plain completions; co-locate Hermes with the Supabase region; keep SSE streaming end-to-end.
- Versioning β freeze the
/v1contract; ~18 callers depend on it.
