In freemansoft/Flutter-AdaptiveCards a demonstration Dart chat server hands a question to a local Ollama model. It asks for the answer as Adaptive Card JSON, a strict, closed-vocabulary schema that a Flutter client renders as interactive UI rather than as text. A set of probes in that repository puts identical questions to different local models. Probe context_fill_probe.dart produced every figure in this article. It runs 25 test cases, one question each. A test case passes only if the reply used an element type that would answer the question. Every other probe in that set had been asking into a nearly empty window. That window holds a system prompt, one question, and at most a short canned exchange showing the model the shape of a card. The literature documents that a long context degrades model behavior. What it does to a model's adherence to an output format is measured less often. The test probe simulates a long conversation by filling three q...
I do a lot of my development and configuration via ssh into my Raspberry Pi Zero over the RNDIS connection. Some models of the Raspberry PIs can be configured with gadget drivers that let the Raspberry pi emulate different devices when plugged into computers via USB. My favorite gadget is the network profile that makes a Raspberry Pi look like an RNDIS-attached network device. All types of network services travel over an RNDIS device without knowing it is a USB hardware connection. A Raspberry Pi shows up as a Remote NDIS (RNDIS) device when you plug the Pi into a PC or Mac via a USB cable. The gadget in the Windows Device Manager picture shows this RNDIS Gadget connectivity between a Windows machine and a Raspberry Pi. The Problem Windows 11 and Windows 10 no longer auto-installs the RNDIS driver that makes magic happen. Windows recognizes that the Raspberry Pi is some type of generic USB COM device. Manually running W indows Update or Upd...
In freemansoft/Flutter-AdaptiveCards a demonstration Dart chat server asks a local Ollama model for an answer as Adaptive Card JSON. That is a strict, closed-vocabulary schema, which a Flutter app renders as interactive UI. The card system prompt that describes the element vocabulary is about 3,755 tokens. The chat server replays up to ten prior exchanges on every turn. Each Ollama runner keeps a prefix cache, so a request whose opening tokens match an earlier one does not have to process them again. Ollama 0.33.3 and later report, in prompt_eval_cached_count , how many of a reply's prompt tokens came from that cache. A chat server can read from that field how much of each prompt the cache served and where it missed. Every reading below comes from ModelBehavior.md , the lab notebook in that repository. A new conversation reuses a cached system prompt on some models and not on others. The split follows how the model stores its context. A pr...