Vancetope — Follow-Up Service
REST endpoint that generates context-aware follow-up suggestions from a text fragment plus cursor position. Single-shot, no Process-Spawn — built on the LightLlmService with Recipe
follow-upas a config profile.Consumed by the chat prompt and text editor in the Web UI (see web-ui); other client surfaces (Mobile, future editors) can use the same endpoint.
See also: light-llm-service recipes how-do-i
1. Purpose & Scope
Problem. Multiple UI surfaces want to display context-aware “what-could-come-next” suggestions:
- Chat Prompt (Edit Mode): User starts typing, Vancetope suggests follow-up words/sentence completions at the cursor.
- Chat Reply (Reply Mode): User looks at the current conversation history, Vancetope suggests a suitable next message. The context can include multiple participants and non-alternating roles.
- Text Editor (Edit Mode): User writes a document, Vancetope suggests follow-up sentences or questions at the cursor.
- Future Surfaces: Inbox Quick-Replies, Wizard Form Free-Text fields, Voice Mode hints.
Implementing each of these use cases as a separate Engine would be overkill: no lifecycle, no history, no tools — just a single LLM call with text context in, suggestion list out.
Solution. A thin FollowUpService in vance-brain that calls the follow-up Recipe via LightLlmService. REST endpoint POST /brain/{tenant}/follow-up/{project}. The service has two structural modes, controlled by the presence or absence of cursor:
- Edit Mode (
cursor != null): Service splitstextat the cursor position intotextBefore/textAfterand sends both to the Prompt. - Reply Mode (
cursor == null): Service sends the entiretextasprecedingContextto the Prompt.
An empty result (no suggestions) is a valid response.
What it is not:
- Not a dedicated Engine — the entire flow is single-shot. Should FollowUp later become multi-stage (e.g., first RAG, then Generate, then Re-Rank), it can be upgraded to a
mark-iiEngine — the Recipe name remains stable. - No streaming — a complete response is returned.
- No Tool-Use — the LLM is not allowed to call Tools in this call.
- No Chat History entry — Follow-Up is a UI aid, not a conversation-relevant action.
2. REST Surface
POST /brain/{tenant}/follow-up/{project}
Authorization: Bearer <jwt>
Content-Type: application/json
Request — Edit Mode (User editing text with cursor):
{
"text": "We should split the migration path ",
"cursor": 31,
"count": 3,
"mode": "chat-prompt"
}
Request — Reply Mode (User replying to preceding text):
{
"text": "Hi, how can I help you today?",
"count": 1,
"mode": "chat-reply"
}
Response (200 OK):
{
"suggestions": [
{ "text": "into multiple phases.", "kind": "completion" },
{ "text": "all at once — what are the risks?", "kind": "continuation" },
{ "text": "coordinate with the Ops team.", "kind": "completion" }
]
}
Path Parameters:
tenant— Tenant slug. Checked by the upstream filter against the JWT claimtid.project— Project name within the Tenant. Auth enforcement:Resource.Project(tenant, project)withAction.READ. Conventionally:_tenantas the default Project if no Project reference is desired (see architektur-scopes-clients).
Body Fields:
| Field | Type | Required | Description |
|---|---|---|---|
text |
string | ✓ | In Edit Mode: complete text the user is editing. In Reply Mode: bounded transcript of the preceding conversation context with role and, if available, speaker markers. Multiple participants and non-alternating roles are allowed. An empty string is allowed in Edit Mode, but meaningless in Reply Mode. |
cursor |
int | ✗ | If set: Edit Mode. Cursor position as a character offset from the beginning. Must be within [0, text.length()] — the service automatically clamps, the controller validates beforehand and returns 400 BAD REQUEST if out-of-range. If omitted/null: Reply Mode. The entire text is considered precedingContext. |
count |
int | ✓ | Desired number of suggestions (max). Server cap at FollowUpService.MAX_COUNT (=10). <1 is rejected with 400. |
mode |
string | ✗ | Free-form hint indicating which UI surface the call originates from (e.g., "chat-prompt", "chat-reply", "text-editor"). Passed as a Pebble variable to the Prompt — orthogonal to the Edit/Reply branch. Reserved for later specialization without contract change. |
Response Schema:
| Field | Type | Description |
|---|---|---|
suggestions |
FollowUpSuggestionDto[] |
List of suggestions, capped to the (clamped) count. Never null. An empty array is allowed. |
FollowUpSuggestionDto:
| Field | Type | Required | Description |
|---|---|---|---|
text |
string | ✓ | The suggested follow-up text. Never empty (parser filters blank). |
kind |
string | ✗ | Optional classification, set by the LLM. Conventional labels: question, clarification, completion, continuation. UI may ignore. |
Error Behavior:
| Status | Condition |
|---|---|
200 + empty suggestions |
LLM found no meaningful follow-up suggestions — not an error, desired behavior. |
| 400 | text null, cursor set but outside [0, text.length()], count < 1. |
| 401 / 403 | JWT missing/invalid or Project READ denied. |
| 500 | LightLlmService throws LightLlmException (Recipe missing, Provider exhausted, Schema-Loop fails). |
3. Architecture
HTTP POST
│
▼
FollowUpController
│ (Auth.enforce: Resource.Project READ)
│ (Validation: text, cursor-bounds if set, count >= 1)
▼
FollowUpService.suggest(text, cursor?, count, mode, tenantId, projectId)
│ - count clamp → [1, MAX_COUNT]
│ - cursor clamp → [0, text.length()] (Edit Mode) or -1 (Reply)
│ - cacheKey = sha256(tenant|project|text|cursor|count|mode)
│ - Caffeine.getIfPresent(cacheKey) → metric outcome=hit, return
│ - cache miss → metric outcome=miss, continue
│ - cursor != null → Edit Mode: split text → textBefore, textAfter
│ - cursor == null → Reply Mode: precedingContext = text
│ - build pebbleVars
▼
LightLlmService.callForJson(LightLlmRequest{ recipeName="follow-up", ... })
│ - Recipe "follow-up" is loaded from Cascade (project → tenant → bundled)
│ - Pebble-renders promptPrefix with {% if precedingContext %} branch
│ - ChatModel.chat() → JSON-Parse → Schema-Validate-Loop
▼
parseSuggestions(rawJson, limit)
│ - tolerant to drift: bare-string-array, missing keys, blank text
│ - drop unusable entries, truncate to limit
│ - cache.put(cacheKey, parsedList)
▼
List<FollowUpSuggestionDto>
│
▼
FollowUpResponseDto { suggestions } → JSON → HTTP 200
Edit Mode (cursor != null): textBefore = text.substring(0, cursor), textAfter = text.substring(cursor). Both are rendered as Pebble variables {{ textBefore }} and {{ textAfter }} into the System Prompt. The LLM thus knows the user’s current position and can produce follow-up suggestions at that exact location.
Reply Mode (cursor == null): The entire text is passed as {{ precedingContext }} to the Prompt. The LLM responds with suggestions on how the user could react to this context (follow-up question, clarification, acknowledgment). The Pebble template branches with {% if precedingContext %}reply-prompt{% else %}edit-prompt{% endif %}.
This branching is semantically necessary: the two tasks (“what comes after the cursor at this point” vs. “how to react to this text”) are semantically different. mode remains as an additional free-form hint for UI surface specialization — orthogonal to the Edit/Reply branch.
FIM Alternative Path: Edit Mode can run via a Completion path instead of this Chat Recipe if ai.alias.default.fim is set — see §5. The flow above describes the Chat path; the cache applies to both paths unchanged.
Result Cache (Caffeine). The service maintains an in-memory LRU cache that covers repeated identical requests:
- Library: Caffeine 3.x — already transitive via Spring Boot 4, explicitly declared as a dependency in
vance-brain/pom.xml. - Key: SHA-256 hex digest over
tenantId | projectId | text | cursor | count | mode(null separator in between). Compact in the map, does not leak user content into string interns; Reply vs. Edit mode is distinguished bycursor = -1vs.>= 0. - Value: parsed
List<FollowUpSuggestionDto>— we cache the typed result, not the raw JSON string, so parsing runs only once. - Eviction:
maximumSize = 500(LRU) +expireAfterWrite = 30 minTTL. - Metric:
vance.followup.cacheCounter with tagoutcome∈ {hit,miss}. Hit-ratio is a direct statement about cache effectiveness. - What it covers: Page reload, multi-tab of the same user, multiple users on a shared Hub Project with identical conversation context, and quick re-displays due to focus toggle in the Composer.
- What it does not cover: Multi-Pod sharing (each Pod has its own cache) — with N Pods, the first call of each Pod instance will still see an LLM miss. Persistent cache via Mongo is v2 if load data shows the need.
- Consistency under Recipe Reload: no explicit invalidation — old cache entries expire after a maximum of 30 min. Acceptable because Recipe edits are not hard cuts anyway, and the effect propagates naturally. If force-invalidate is needed later, an
@EventListenerwill attach to the Recipe reload event and callcache.invalidateAll().
Tolerant Parsing Strategy: The service accepts several drift modes of the LLM response:
- Object Form:
[{ "text": "...", "kind": "..." }, ...]— the canonical form. - Bare-String Form:
["...", "..."]— occasionally chosen by the LLM, interpreted as{ text, kind: null }. - Missing
suggestionskey → empty list. - Items without
textor with blanktext→ dropped.
This keeps the service robust without bloating the Prompt with excessive format reminders.
Empty-Result as First-Class-Outcome: The service explicitly returns 200 + [], no 204 and no error. This simplifies client code (no status check needed) and respects reality: not every cursor context warrants suggestions.
4. Recipe follow-up
vance-brain/src/main/resources/vance-defaults/_vance/recipes/follow-up.yaml:
description: |
Internal config profile for the LightLlmService consumed by
FollowUpService (backend of the follow-up suggestion REST endpoint).
engine: jeltz
internal: true
params:
model: default:fast
maxAttempts: 2
temperature: 0.7
promptPrefix: |
You are Vancetope's follow-up suggestion component. Produce up to
{{ count }} short follow-up suggestions matching the context below.
{% if precedingContext %}
## Preceding text (the user is responding to this)
{{ precedingContext }}
## Task
Suggest up to {{ count }} short ways the user could **react** —
follow-up question, clarification, acknowledgement.
{% else %}
## Text before cursor
{% if textBefore %}{{ textBefore }}{% else %}(empty){% endif %}
## Text after cursor
{% if textAfter %}{{ textAfter }}{% else %}(empty){% endif %}
## Task
Suggest up to {{ count }} short ways the user could **continue**
typing at the cursor position.
{% endif %}
## Reply format
Respond with exactly ONE JSON object — { "suggestions": [...] } — no
prose, no markdown code fences.
tags:
- internal
- followup
Engine Choice: jeltz provides the schema validation loop “for free”. If the LLM does not provide parseable JSON, the call is retried with a correction hint (up to maxAttempts). Error case: SchemaValidationException → 500.
Pebble Branching: {% if precedingContext %} switches between Reply and Edit Prompt. This means the two-mode logic is fully declarative in the Recipe; the service only decides which Pebble vars to set.
Temperature 0.7 — intentionally higher than for Discovery (0.0), because creativity is needed here, not deterministic selection from a catalog.
Model default:fast — Follow-Up is latency-sensitive (user waits in the UI). The inexpensive/fast model per Tenant is resolved via alias cascade (see llm-resource-management §3a).
Tenant Override: Tenants can place their own recipes/follow-up.yaml via the Recipe Cascade under _vance/recipes/follow-up.yaml. internal: true is retained — otherwise, the Recipe would collide with the standard Recipe selector.
5. FIM Completion Path (Edit Mode, Optional)
Edit Mode can use a Completion Model (Fill-In-the-Middle) instead of a Chat Model — models like Qwen2.5/Qwen3-Coder, DeepSeek-Coder, Codestral, or StarCoder, which are trained to fill a gap between prefix and suffix using their own FIM control tokens. Reply Mode never uses FIM — there is no suffix to fill against.
5.1 Gate and Model Selection — ai.alias.default.fim
A setting decides the path (Tenant/Project cascade like all ai.alias.*):
ai.alias.default.fim = lmstudio:qwen3-coder-30b
- Unset → both modes run via the Chat path (Recipe
follow-upvia LightLlm) — the existing behavior, unchanged. - Set → Edit Mode calls run via the FIM path; the value is a normal Model Spec (instance, provider, or another alias) and goes through the
AiModelResolver.
Operators set the alias in the “LLM Settings” form (llm-setup, field aliasFim); the picker uses the ai-fim-models choices source and lists only models with a resolved fimTemplate — otherwise, a Chat Model would be a fail-closed error at runtime. The LLM-side Manual for creating Model Docs (manual_read('ai-model-catalog')) documents the fimTemplate field, including a family table.
5.2 Call Shape — fimTemplate Quirk (Per Model, Not Global)
The FIM token convention is family-specific, so the Prompt shape is per-model metadata and not a global constant. The fimTemplate field follows the same two-layer resolution as messageParser (see llm-resource-management §3a/§4.1.1): explicit per-model YAML wins, family patterns from the bundled model-quirks.yaml fill the gap:
# model-quirks.yaml (bundled Patterns, excerpt)
rules:
- match: "qwen*coder*"
fimTemplate: "<fim_prefix>{prefix}<fim_suffix>{suffix}<fim_middle>"
- match: "deepseek-coder*"
fimTemplate: "<|fim▁begin|>{prefix}<|fim▁hole|>{suffix}<|fim▁end|>"
- match: "codestral*"
fimTemplate: "[PREFIX]{prefix}[SUFFIX]{suffix}[MIDDLE]"
The markers {prefix} and {suffix} are spliced with textBefore/textAfter — the template itself is parsed, the user content never is (a literal {suffix} in the edited text remains untouched).
Fail-closed: Alias set, but the resolved model has no fimTemplate → error with a clear message (“Setting points to a model without FIM shape”), no silent fallback to the Chat path. Such a fallback would mask a configuration error and replace it with double costs. The separation is deliberate: the setting decides which model, the quirk field decides how to query.
5.3 Architecture
FollowUpService (Edit Mode, Cache Miss)
│ ai.alias.default.fim set?
│ no → Chat Path (§3, unchanged)
│ yes:
▼
FimCompletionService (brain.ai.fim)
│ - Resolve spec (AiModelResolver) → AiChatConfig
│ - ModelCatalog.lookupOrDefault → ModelInfo.fimTemplate
│ - Render template (splice prefix/suffix)
│ - Single User Message over the standard Chat plate
│ (ChatBehaviorBuilder.resolveOne → AiModelService.createChat)
▼
Raw Continuation → Trim end markers (FIM end tokens) → strip
▼
1 × FollowUpSuggestionDto { text, kind: "completion" }
Transport over the Chat Wire: vLLM, llama.cpp, and LM Studio — where these models typically run — accept the FIM token sequence as a simple User Message on the OpenAI-compatible Chat Endpoint. This allows the entire existing plate (alias resolution, provider adapter, usage ledger, audit) to be used; a dedicated /v1/completions-with-suffix transport will only become a provider feature if a specific endpoint rejects the Chat form.
Fixed, conservative Call Params (deliberately no Recipe — the FIM path has no Pebble Prompt, so there is no config profile; if needed, later as settings): Temperature 0.2, maxTokens 256, Stop sequences = FIM end markers, Sync deadline 30s, no System Prompt, no JSON schema loop.
5.4 Result Semantics
- FIM delivers exactly one Continuation;
countis ignored on this path (documented, not hidden). To getcount > 1diverse suggestions, leave the alias unset. - Blank-Middle (after trim) is a valid “nothing goes here” →
200 + empty list, no error, no silent Chat fallback. - The REST contract does not change — client-side, the path is invisible; the suggestion carries
kind: "completion".
5.5 Metrics & Audit
vance.fim.calls{outcome=success|blank|error, caller}andvance.fim.duration{caller}— Caller is the consuming service (e.g.,follow-up-fim), low-cardinality.- Audit like Light Calls:
llmLightCallwith Caller as Recipe name; the Usage Ledger is booked by the Accounting Decorator in the Provider, as with any Chat call. - The FollowUp cache remains unchanged: the cache key already distinguishes Edit/Reply via
cursor, and the path decision depends only on settings, not on request content.
6. DTO Contracts
vance-api/de/mhus/vance/api/followup/ — @GenerateTypeScript("followup") annotated, TS output to client/packages/generated/src/followup/:
FollowUpRequestDto— Request body (see §2).FollowUpResponseDto— Response.FollowUpSuggestionDto— Single suggestion.
Validation annotations (@NotNull, @Min) are enforced by the Controller via @Valid; edge cases (cursor-out-of-range vs. text-length) are additionally checked manually by the Controller, as this is a cross-field constraint.
7. Client Integration
Web (vance-face):
import { FollowUpRequestDto, FollowUpResponseDto } from '@vance/generated';
import { restPost } from '@vance/shared/rest';
// Edit mode — user is typing in the chat prompt
const editResp: FollowUpResponseDto = await restPost(
`/brain/${tenant}/follow-up/${project ?? '_tenant'}`,
{ text, cursor, count: 3, mode: 'chat-prompt' } satisfies FollowUpRequestDto,
);
// Reply mode — input is empty; send a bounded, speaker-aware
// transcript of the recent main conversation
const replyResp: FollowUpResponseDto = await restPost(
`/brain/${tenant}/follow-up/${project ?? '_tenant'}`,
{ text: recentConversationTranscript, count: 1, mode: 'chat-reply' } satisfies FollowUpRequestDto,
);
Chat Editor “Ghost Bubble” (Reply Mode, v1):
- Trigger: new relevant chat turn fully persisted and input field empty and composer focused (fetch gate).
- Display: Ghost bubble below the last context message, muted color + italic, with hint icon (e.g.,
↹ Space). - Adoption: User presses Space → suggestion is written into the input field (plus appended space, shell autosuggestion style). Alternatively Tab → without space. Click on bubble → also adopt.
- Hiding: As soon as
input.length > 0and the first character is not a space, the bubble disappears. - Reappearance: Input back to empty → bubble re-appears (cached, no new call).
- Invalidation: new relevant chat turn → old suggestion discarded, new call with the updated transcript.
- Fetch Gate (Strategy B): The REST call runs only when the textarea is focused. The client composable caches per
(projectId, anchorMessageId, transcript), and re-focusing does not lead to a second call. This way, we don’t pay for an LLM call for users who don’t want to reply.
Markdown Document Editor “Tooltip Suggestion” (Edit Mode, v1):
- Implementation:
followUpExtensionfrom@vance/componentsas a CodeMirror extension. Active only if<CodeEditor :follow-up="…" />is set; inDocumentApp.vue, this is only the case for Markdown documents in the Edit tab. - Trigger: User presses
Ctrl+.(Mac:Cmd+.) in the editor. On-demand, never automatically while typing — respects that the user decides when they want help. (Ctrl/Cmd+Spacewould be more intuitive, but is reserved on macOS by the OS for Spotlight or the IME switcher;Mod-.is platform-neutral free.) - Display: CodeMirror
showTooltipdirectly under the cursor, muted color + italic, with hint label (↹ Tab). - Adoption: Tab → suggestion is inserted at the cursor, cursor jumps to the end of the insertion, tooltip disappears.
- Discarding: Escape → tooltip disappears without insertion. Any document change or cursor movement also discards the suggestion (anchor position would then be stale).
- Stale-Drop: Pending fetch is discarded via sequence counter if a newer trigger or document change occurs.
- Other Editors (Code with Java/Python/YAML, Office Editor, Setting Forms): deliberately no Follow-Up — code completion is the task of the IDE, not Vancetope. Only Markdown.
Other UI Surfaces (Inbox Quick-Reply, Wizards, Voice): no Follow-Up in v1.
Edit Mode in the Chat Prompt is not v1.
Mobile (vance-fingers): identical endpoint, same types from @vance/generated. Mode can set, e.g., "chat-prompt" or "voice-input".
CLI (vance-foot): Reply-mode follow-up loads a limited history tail on idle trigger, serializes roles and USER speaker names, and displays the suggestion as REPL ghost text. The cursor-based Edit Mode remains unused.
8. What v1 Does Not Do
- No RAG. Suggestions are based solely on the provided text context. As soon as Project Memory is to be included, it’s v2: either pre-retrieval (embedding lookup on
textBeforeenvironment) or upgrade to a realmark-iiEngine with a Tool-Call loop. This breaks the single-shot model — therefore deliberately postponed. - No persistent caching. The service maintains an in-memory Caffeine cache (LRU, 500 entries, 30 min TTL, cache key = SHA-256 over
tenantId + projectId + text + cursor + count + mode), which covers page reloads, multi-tab, and shared Hubs. Multi-Pod setups lose cross-Pod sharing — Mongo persistence is v2. Hit/Miss count undervance.followup.cache{outcome=hit|miss}. - No Persona adaptation. The Prompt is generic; persona-specific suggestions (e.g., specialized language tone) are v2.
- No streaming. Suggestions appear as a block once the LLM call returns. For
default:fast(~500ms p50), this is acceptable. - No special rate limit handling. Standard quota cascade (see llm-resource-management) applies — excessive calls are blocked like normal LLM calls.
- No diverse FIM suggestions. The FIM path (§5) delivers exactly one Continuation per call — N diverse suggestions via N temperature-sampled calls would be possible but increase call cost without guaranteeing diversity.
- No dedicated Completion Endpoint. FIM transport runs over the Chat wire (§5.3);
/v1/completionswith asuffixparameter will only become a provider feature if a specific endpoint rejects the Chat form.
9. Tests
vance-brain/src/test/java/de/mhus/vance/brain/followup/FollowUpServiceTest.java:
- Cursor split (beginning, middle, end, out-of-range).
- Count clamping (MAX_COUNT) and truncation (LLM ignores limit).
- Parser tolerance: object-form, bare-strings, missing-key, blank-text, empty array.
- Mode passthrough (set vs. null/blank → omitted).
- LightLlmService wiring (Recipe name, Tenant/Project Scope, Schema set).
- FIM path selection: configured → 1 Completion without LightLlm call; blank → empty list; Reply Mode never FIM; unconfigured → Chat path; cache also applies to FIM.
- Validation (null text, blank tenant).
vance-brain/src/test/java/de/mhus/vance/brain/ai/fim/FimCompletionServiceTest.java:
- Gate (
isConfigured), not-configured error, fail-closed on missingfimTemplate(no Provider call). - Template splicing (markers only in template, never in user content), malformed-template rejection.
- End-marker trim, blank-Middle, Provider error →
FimException+ metric outcomes.
vance-brain/src/test/java/de/mhus/vance/brain/ai/ModelQuirksTest.java: FIM family patterns, per-field independence of fimTemplate, malformed-rule drop.
LightLlmService and FimCompletionService mocked — no actual LLM calls in unit tests. End-to-end is tested via qa/ai-test/ opt-in (see CLAUDE.md section “Tests”) when the use case is included in QA.
10. Extension Paths
- MarkII Engine — If FollowUp becomes multi-stage (RAG → Generate → Re-Rank) or needs Tool-Use (e.g., code lookup in Project Files),
mark-iiwill be implemented as a real Engine. The REST contract remains stable; the Recipefollow-up.yamlswitches toengine: mark-ii. - Per-Mode Recipes — If Chat Prompt and Text Editor require very different suggestions, the service can be extended to
follow-up-${mode}Recipe lookup with a fallback tofollow-up. - Manuals Hook — If the model needs to access predefined Domain Manuals, this works like Discovery:
SourceCatalogBuilderfor FollowUp-specific Manuals, appended to Pebble vars. - Inline Ghost Text in Editor — The FIM path (§5) makes continue-while-typing cheap (small local model, no JSON loop). An inline ghost text in the Workpage Editor would be the natural next consumer of
FimCompletionService— today’s Ctrl/Cmd+.-tooltip remains an on-demand Chat path. This would be a client feature with its ownmode, no contract change.