Skip to content

Hermes as the AI gateway β€” "catch everything (AI)" ​

Centralize all LLM generation behind Hermes so model selection, knowledge/RAG, logging, rate-limits, the shadow+eval gate, and data-residency are enforced in one place. Scope note: canonical source lives here under services/hermes-server/ and supabase/functions/_shared/aiClient.ts. The Flutter client is gateway-agnostic (it streams the OpenAI-SSE shape). Pairs with HERMES_MODEL_CONFIG.md.

Principle ​

Edge functions stay the entry points (stable contracts, auth, data fetching). Their LLM step delegates to Hermes. Hermes is the AI brain, not a proxy for all traffic.

app / cron ─▢ edge fn (auth + data fetch) ─▢ Hermes gateway ─▢ provider (CF AI Gateway ─▢ Gemini/…)
                                              β”‚ model pin (hermes_config)
                                              β”‚ RAG (hermes_embeddings)
                                              β”‚ log (ai_usage_logs) + rate-limit (ai_usage_budgets)
                                              β”‚ shadow/eval (vs model.*.target)
                                              └─ (Hermes down?) edge fn falls back to direct gateway

In scope (route through Hermes) vs not ​

AI/LLM functions (~18 β€” route): ai-assistant, ai-regression, ai-valuation-analysis, ai-valuation-report, ai-data-analyst, ai-market-intelligence, ai-marketing-caption, ai-property-search, generate-market-report, generate-property-image, whatsapp-ai-chatbot, whatsapp-summarize, smart-whatsapp-alerts, suggest-price, suggest-parameters, classify-document, scan-document, research-rental-rates.

NOT in scope (transport / data / auth β€” leave as-is): mailbox-api, gmail-import-*, send-email, paci-*, import-*, all whatsapp-* webhooks/senders, process-payment, webauthn-*, push, cron reminders, etc.

The enabler ​

Most AI functions already share _shared/aiClient.ts (callAI). Pointing that one helper at Hermes migrates most of them at once. Exceptions that bypass the shared helper and need individual work:

  • generate-market-report β€” calls Gemini directly with tool-calling.
  • ai-data-analyst, ai-market-intelligence β€” streaming (need Hermes SSE).

Gateway contract (Flutter-owned Hermes service) ​

POST /v1/completions            # non-streaming, OpenAI-compatible
POST /v1/completions/stream     # SSE: data: {choices:[{delta:{content}}]}
Authorization: Bearer <hermes service key>   # hermes_api_keys (service-to-service)

body: {
  messages: [{ role, content }],
  role: "chat" | "reasoning",          # -> hermes_config model.chat / model.reasoning
  response_format?: "json_object" | "text",
  tools?, tool_choice?,                # passthrough for generate-market-report
  temperature?, max_tokens?,
  context_hint?: string,               # opt-in RAG selector (use_case / entity)
  metadata: { function_name, use_case?, user_id? }   # for logging + rate-limit
}
returns: { choices:[{ message:{ content } }], model, usage:{ prompt_tokens, completion_tokens, total_tokens } }

Per call, Hermes:

  1. resolves model = hermes_config["model." + role] (the central pin);
  2. if a model.<role>.target shadow is enabled β†’ also calls it, writes a side-by-side row (reuse the whatsapp_shadow_replies pattern);
  3. optionally injects RAG context from hermes_embeddings (by context_hint/use_case);
  4. enforces rate-limit + ai_usage_budgets keyed on metadata.user_id;
  5. calls the provider via the Cloudflare AI Gateway with model;
  6. logs ai_usage_logs { model, function_name, user_id, prompt/completion/ total tokens, estimated_cost_usd, duration_ms, request_id, use_case };
  7. returns the OpenAI-compatible shape.

The callAI β†’ Hermes shim (_shared/aiClient.ts) ​

ts
import { getModel } from "./modelConfig.ts"; // hermes_config reader

const HERMES = Deno.env.get("HERMES_URL") ?? "https://hermes.aldilaijan.com/api";
const HERMES_KEY = Deno.env.get("HERMES_SERVICE_KEY")!;
const GATEWAY_ON = (Deno.env.get("AI_GATEWAY_ENABLED") ?? "true") === "true";

export async function callAI(o: {
  messages: { role: string; content: string }[];
  role?: "chat" | "reasoning";
  responseFormat?: "json_object" | "text";
  temperature?: number;
  maxTokens?: number;
  meta: { functionName: string; useCase?: string; userId?: string };
}): Promise<{ content: string; model: string }> {
  if (GATEWAY_ON) {
    try {
      const r = await fetch(`${HERMES}/v1/completions`, {
        method: "POST",
        headers: { "content-type": "application/json", authorization: `Bearer ${HERMES_KEY}` },
        body: JSON.stringify({
          messages: o.messages,
          role: o.role ?? "chat",
          response_format: o.responseFormat,
          temperature: o.temperature,
          max_tokens: o.maxTokens,
          metadata: o.meta,
        }),
        signal: AbortSignal.timeout(45_000),
      });
      if (r.ok) {
        const j = await r.json();
        return { content: j.choices[0].message.content, model: j.model };
      }
      console.warn(`callAI: Hermes ${r.status} β€” falling back to direct`);
    } catch (e) {
      console.warn("callAI: Hermes unreachable β€” falling back to direct:", e);
    }
  }
  // Fallback: existing direct provider path, using the SAME pinned model so
  // behaviour matches. (Keep the current callProviderDirect implementation.)
  return callProviderDirect({ ...o, model: await getModel(o.role ?? "chat") });
}

A streamAI twin does the same against /v1/completions/stream, falling back to the current direct-SSE path. Toggling AI_GATEWAY_ENABLED=false disables routing instantly (kill-switch) β€” also mirror it as a hermes_config key (ai.gateway.enabled) so it's flippable without redeploy.

Phased migration ​

  1. Stand up /v1/completions + /stream in Hermes: read hermes_config, call the CF AI Gateway, write ai_usage_logs. Add HERMES_SERVICE_KEY.
  2. Pilot one function — ai-valuation-analysis — through callAI→Hermes. Shadow vs direct; compare quality / cost / latency from ai_usage_logs.
  3. Flip the shared callAI β†’ Hermes (most functions migrate at once). Keep the fallback on; watch ai_usage_logs.error_message + p95 latency.
  4. Exceptions: refactor generate-market-report (tool-calling passthrough) and the streaming ai-data-analyst / ai-market-intelligence onto /stream.
  5. Turn on RAG-context injection + the shadow/eval gate uniformly.

Rollback at every step: AI_GATEWAY_ENABLED=false (or the hermes_config key) reverts to direct calls in one flip; the per-call fallback already degrades gracefully on a Hermes blip.

Why it's worth it ​

  • One model pin (hermes_config) actually governs all AI, not just chat.
  • One ai_usage_logs stream β†’ one cost/latency/quality dashboard.
  • The knowledge/RAG layer + house voice apply to every AI surface.
  • One place to enforce data-residency (what PII β€” incl. anything touching registry_voters β€” reaches which provider).

Trade-offs (design for these) ​

  • SPOF β†’ Hermes HA + the per-call fallback above (degrade, don't die).
  • Latency β†’ thin pass-through for plain completions; co-locate Hermes with the Supabase region; keep SSE streaming end-to-end.
  • Versioning β†’ freeze the /v1 contract; ~18 callers depend on it.

Aldilaijan & Khobara Real Estate Platform