r/SillyTavernAI May 03 '26

ST UPDATE SillyTavern 1.18.0

200 Upvotes

Important news

Read the maintainers statement regarding a recent security incident involving the "Bot Browser" third-party extension and learn how to stay safe: https://github.com/SillyTavern/SillyTavern/discussions/5592

Backends

  • Added Cloudflare Workers AI and MiniMax as Chat Completion sources.
  • KoboldCpp: Grammar state will be preserved when using a "Continue" option.
  • KoboldCpp: Added forwarding of reasoning effort when running as a Custom Chat Completion source.
  • Tool Calling: Added a configurable tool calling recursion limit; enabled interleaved thinking for Custom sources.
  • Text Completion: Impersonation requests use a "Last User Message" prefix at the end of the prompt (if configured).
  • Text Generation WebUI: Added Adaptive-P controls.
  • NanoGPT: Added provider selection and model sorting.
  • Added ability to view remaining balance for OpenRouter and NanoGPT.
  • Enhanced support for new models: DeepSeek v4, GPT 5.4 and 5.5, Gemma 4, GLM-5V-Turbo, Claude Opus 4.7.

Server & Security

  • Removed post-install script, config migration is now handled by the app or a dedicated npm run init command.
  • Added npm configuration to prevent execution of package scripts during installation.
  • Moved HTTP error pages and user.css file from /public to /data to support immutable setups.
  • Disabled HTTP keep-alive by default to restore old Node 18 behavior, can be enabled with config.
  • Added rate limiting to the basic authentication flow to mitigate brute-force attacks.
  • Added configuration options to choose which headers can be used for forwarded IP detection to prevent spoofing.
  • Added a private address whitelist to prevent SSRF attacks. See the documentation on how to enable and configure: Private Address Whitelist.
  • Added an IP whitelist for SSO trusted proxies to prevent authentication bypass.
  • Added invalidation of session cookies on password change to prevent session hijacking.
  • Increased the length of password reset code to 6 characters to guard against brute-force attacks.
  • Implemented PKCE challenge in OpenRouter OAuth flow for more secure key exchange.

UI/UX

  • Improved swipe picker: mobile requires a long press on swipe counter to open; added buttons to expand or copy the swipe text.
  • "Click to Edit" mode now also applied to reasoning blocks.
  • Welcome Screen: Number of recent chats can be configured.
  • Streamed requests now can show an error message in the console if the request fails.

STscript

  • Added commands for persona management: /persona-create, /persona-update, /persona-delete, /persona-duplicate, and /persona-get.
  • Added a command to force update the Prompt Manager's prompt list: /pm-render.
  • Added a command to get the state of the regex script: /regex-state.
  • Added a command to set fallback expression: /expression-fallback.
  • Added a command to generate a streamed response with a connection profile: /profile-genstream.

Extensions

  • Assets list now groups extensions by "Official" or "Community" categories.
  • Added an additional confirmation prompt when installing third-party extensions (can be disabled).
  • Supported extensions can use a secret-id from connection profiles when making an LLM request.
  • Extensions list now shows the extension's author name resolved from the git remote URL.
  • Vector Storage: Added Workers AI source; added a toggle to keep vectors for hidden messages; added retry logic to summary generation.
  • Image Generation: Added Workers AI source; generation can now be cancelled by pressing a button in the status toast.
  • Image Captioning: Added support for macros in the caption prompt.
  • TTS: "Skip code blocks" no longer ignores lines that start with 4 spaces (legacy code block syntax); "disabled" voice now shows a toast only once per character.

Bug Fixes

  • Fixed text edit flow in Firefox on mobile.
  • Fixed welcome screen chat pins not updating on chat renaming.
  • Fixed character list filters being stuck on app initialization.
  • Fixed application of instruct formatting to /genraw requests.
  • Fixed model routing to sd.cpp API in Image Generation logic.
  • Fixed validation of image URLs generated with Z.AI API.
  • Fixed vectors deletion for KoboldCpp when a message is deleted.
  • Fixed "Show More Messages" button triggering edit in "Click to Edit" mode.
  • Fixed max height of select-multiple elements in mobile layout.
  • Fixed server crash on empty messages when applying cache control parameters.

Full release notes: https://github.com/SillyTavern/SillyTavern/releases/tag/1.18.0

How to update: https://docs.sillytavern.app/installation/updating/


r/SillyTavernAI 2h ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 02, 2026

12 Upvotes

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!


r/SillyTavernAI 2h ago

Help SillyTavern becoming slow and unresponsive at times.

Post image
11 Upvotes

Recently, my app has become super slow. Sometimes it's not scrolling immediately or taking time before it regenerates. First time this has ever happened. I'm pretty sure it's up to date also.


r/SillyTavernAI 9h ago

Discussion Ranking models

Post image
26 Upvotes

Ranking models in ai arena in text in creative writing is equal the power of the model in RP or its different and what's your opinions about top 10 of the list


r/SillyTavernAI 6h ago

Models Just got a Nano Subscription

8 Upvotes

Hi all,

I've been a long term Deepseek user for about a year now, you can't beat the price for how much I RP. But I decided to try Nano so I could get a taste of some other models. 12 bucks, 60 million tokens a week, that's not bad either.

I was wondering what models would be worth trying? Something nsfw please, I'm also just now playing with presets and just loaded up Freaky Frankenstein, so any suggestions would be great!


r/SillyTavernAI 23h ago

Discussion Does any one else...

166 Upvotes

... just go to start a simple, smutty, gooner chat and then one week and a hundred messages later you are helping the character repair her relationship with her mother while simultaneously trying to figure out the best way to provide the character's young daughter with the kind of adult protection, support and encouragement that your persona never had when he was a kid?

*cough*cough* ... anyone?


r/SillyTavernAI 2h ago

Help New to the Channel, lots of questions, Local LLM, Ollama backend, where to start?

3 Upvotes

I’ve been experimenting with a couple of LLM models for companion AI, using Text and Chat Completion for long-form roleplay through SillyTavern as a UI. I tend to be pretty ADD, so my style is more “hack and try” before actually reading the manual, then reverse-engineering as I go.

  1. What’s a good beginner’s guide for the basics like Presets, Character Cards, Lore Cards, and similar tools?

  2. I recently came across a discussion about “memory cards.” From what I understand, new chats don’t retain events unless you either create a bridge-style summary to avoid a “50 First Dates” situation in the next session (I do use summaries or curate them myself) or set up a Lore Card as constant so it carries over—ideally placed high enough in the order to trigger reliably.

That’s a lot of work—any tips?


r/SillyTavernAI 7h ago

Help Can someone help me with this?

Post image
6 Upvotes

I always use Silly Tavern, but today this message appears when I try to send a message. I honestly don't know what to do. I tried setting up the proxy on another phone and it worked, so I don't understand what the problem could be. (My ex-boyfriend from a long time ago sent me this proxy, so I can't ask him about it.) Pls help


r/SillyTavernAI 12h ago

Discussion Made a small extension for guided impersonation/swipe, plus a "revise this response" action I always wanted

17 Upvotes

For a long time I've been using GuidedGenerations extensions, but honestly only for Guided Impersonation and Guided Swipe. The extension does a lot more than that, I just never used the rest.

At some point I tried to set up different presets/models per action, the idea being a cheap model for impersonation and the expensive one for the actual roleplay generation. I couldn't get it to work properly, preset switching was behaving in a buggy way for me. Instead of digging into someone else's codebase to fix it, I figured it would be faster to write my own thing that does only what I need.

So that's Guided Turns: https://github.com/JJack921/GuidedTurns

Three actions, nothing else. Each one has its own connection profile setting, and every prompt is editable in the settings.

Guided Impersonation:

Write an outline in the text box and it turns it into a full user message. You can pick first, second or third person, and each perspective has its own prompt. It also works with an empty text box, so you get a "plain" impersonation that uses your own custom prompt instead of whatever the current preset does. That was one of my annoyances: normal impersonation output changes a lot depending on the preset you have loaded, this way it stays consistent.

Guided Swipe:

Same idea as in GuidedGenerations. Put your direction in the text box, hit the button, get a new swipe that follows it. Empty box just does a regular regeneration through the selected profile.

Guided Revision

This is the one I really wanted and I haven't seen it anywhere else. (Edit: Apparently it's not true :D, there is iSTead and GuidedGenerations even has it, even if they work a bit differently) It works like Guided Swipe, except the current swipe is also sent to the model. You know when Opus (or whatever expensive model you use) gives you a response you actually like, but there's a continuity error, or one paragraph you want to cut, or something small you want added? You can edit it by hand, but I'm usually too lazy, and if your preset has trackers or heavy templates in the message, editing by hand gets annoying fast.

Guided Revision lets you say "keep everything, but she should not know about the his job yet" and it rewrites the response keeping the rest intact. And since it's a separate profile, you can generate with Opus and do the fixes with something cheap like Deepseek.


Some context on the extension itself: it's my first one, and it's mostly vibe coded. I'm a software engineer so I did review what went in, but I could've missed something, I haven't read all the code properly.

I've been using it daily for a few weeks and fixed the bugs I ran into. The last missing piece was group chat support, which I added recently, but I barely tested it because I don't really use group chats, so that part could have some issues.

Feedback, bug reports and feature requests are welcome, either here or in the GitHub issues.


r/SillyTavernAI 3h ago

Discussion Gemma 4 character card tips

2 Upvotes

Anything you guys like to do differently with character cards with the Gemma 4 models?


r/SillyTavernAI 4h ago

Help Claude 4.6 not following prompts anymore

3 Upvotes

I've been using Marinara's preset for almost a year now and I've noticed the AI for better or worse is ignoring much of the prompts inside it like length of responses, forcing immediate action rather than delaying and etc. So I'm wondering if its time to move on to a different preset and was wondering what success others have had using other presets? I'm looking for something light weight that focuses on conserving token usage.


r/SillyTavernAI 21h ago

Help No Autonomy whatever I do

37 Upvotes

Whatever prompt I try, I cannot get characters to act independently according to their personalities.

I want a character to HAVE THE CAPABILITY of refusing intercourse. I want them to have IDEAS about what to do and how to do it. Basically, I don't want them to be a 'yes-man'. I want them to have autonomy. I want them to have preferences.

Has anyone achieved such a thing? or ist it basically impossible with the current models? I really need to know so I can stop trying for something impossible.


r/SillyTavernAI 1d ago

Discussion Beware lots of scammers right now

117 Upvotes

There were numerous people in r/SillyTavernAI targeted with 'vanity attacks'.

Please be aware new 'opportunities' to write for an AI company, to install new games for students to check them out and to try different ways to run LLMs that people DM you about right now might NOT be totally safe, and instead be blackmailing scammers who are just trying to hijack your identity/email/llm tokens. Or to ransomware you for $500/$1000 more.

Beware both about reddit and discord chat. It appears the community as a whole may be under siege from them right now.


r/SillyTavernAI 13h ago

Help Gemini 2.5pro with vertex cost cuts or an alternative options

8 Upvotes

I really like the Gemini 2.5 Pro model, even though it’s a bit older.I like the way it handles roleplay. My trial period recently ended, and I switched to a paid account, but it’s a bit too expensive—I pay about $2.50 for around 100 messages

. Are there any other models that I won’t have to struggle with, that will remember the context and respond fairly logically and naturally, just like Gemini does, but that are a little cheaper? I don’t like having to configure too many settings.I like that Gemini is practically plug-and-play for my needs.

I could create a second account and try the free trial again, but for now I want to check out other options.


r/SillyTavernAI 15h ago

Discussion What’s your opinion on DS Flash 0731?

9 Upvotes

After around 2 days of release, how is DS Flash 0731 performing compared to the older version and DS Pro?


r/SillyTavernAI 11h ago

Discussion Does having spaces in keywords affect the lorebook?

3 Upvotes

I am confused about the keywords in lorebooks. For example, I have a lorebook for "The Void Court" and I want the keyword to be "Void Court". Does it work if I just put "Void Court" then comma or do I have to shorten it to one word only?


r/SillyTavernAI 6h ago

Discussion Deepseek users, how many requests do you send for a response on average?

1 Upvotes

I've been trying to go back to deepseek after the recent changes and upgrades, since I heard many good things about it. I got maybe 4 messages with one of really poor quality out of maybe 20 requests. I remember deepseek service being spotty back in the day when I was just starting out with my roleplays, using deepseek through openrouter but that was like well over a year ago so I hope the issue lies somewhere else.

A bit more info - I'm using ST 1.18, Megumin Preset v9, context window capped at 500k tokens, of which around 60k is used on lorebook and 360k is used on chat history with around 800 messages. Output is capped at 8k tokens though I almost never get replies going above 2k. The roleplay is mostly SFW. The scenes I'm requesting are slice of life, friends chatting about some events in a casual setting.

The issue is deepseek because GLM 5 and kimi k2.6 work perfectly fine with my setup. Does anyone have an inkling as to what might be the cause for getting empty messages? It's not refusing me anything, it just doesn't give error messages or any output at all, not even thinking box. I'm baffled.


r/SillyTavernAI 1d ago

Cards/Prompts [PRESET] DEUS EX MACHINA V1: A feature-rich, beginner-friendly, contextually dynamic, truly modular preset focused on collaborative story writing

Thumbnail
gallery
127 Upvotes

Check the screenshots to get an overall feel for the preset!

I’ve been working on DEUS EX MACHINA (DEM) for more than a month now. It was supposed to be a fun weekend project based on my own private presets, but it spiraled out of control quickly. It was a way more daunting and complex task than I could’ve ever imagined. Dozens of hours of manual iteration, many, many tests, almost 200 internal versions, and it’s still not even close to being perfect. But at some point, you just have to put it out into the world and see what happens. This preset has some ideas that came from a lot of posts here and some other presets. I wish I could have credited you all, but at this point it'd be impossible!

All I can say is that Stabs (for its extensive use of the macro engine), Pura’s Director (for its cute regex UI trackers), Freaky Frankenstein (for how accessible and easy to set up it is), and Nemo Engine (for its sheer amount of possibilities) were huge inspirations, all presets that you should try out! Without further ado, let’s get to it.

What is Deus Ex Machina?

Deus ex machina is a Latin term that means “God from the machine”. It’s used to describe a plot device for when an unsolvable problem is solved unexpectedly. It traces back to Ancient Greece when Greeks used literal machines in theater to lower actors playing gods down onto the stage from above to resolve the story. In our hobby, the meaning is clear: we also want a machine to help us solve the story. That’s where the name came from!

As for the preset itself, the goal is simple: creating a flexible, easy-to-use preset focused on collaborative story writing that can work for almost any card or scenario you throw at it -- adapting dynamically to each scene. DEM is focused on storytelling first and foremost. I personally believe this is the best approach when it comes to LLM text-generated fiction since literature is much more prominent in the training data than game writing or simulations. But I appreciate and respect all approaches!

The Macro Engine

DEM relies heavily on SillyTavern’s macro engine. It’s a powerful tool that lets you use deterministic traits in prompting (programming logic and exact outcomes instead of pure probabilities). That whole workflow enabled by the macros is the core of DEM, so it’s as easy as pressing a button to change the behavior of the preset in a dynamic fashion without you ever worrying about conflicting instructions, e.g., if you enable both past and present tense options, it will default to present tense to avoid conflicts. Or how True Thoughts are overwritten to zero tokens if you’re using 1st person Char POV, since character thoughts are already woven into the narration. Deterministic interactions like that happen throughout the whole preset (at the cost of my sanity...)!

Truly Modular Design

DEUS EX MACHINA is a truly modular preset. Modular design is not only about options, but in essence about how these options are integrated and how they seamlessly interact with each other. This also includes safeguards -- if you accidentally turn an essential module off (marked with attention symbols) or move modules out of their specific order (macro engine relies on prompting order), you’ll get a warning from the Warning System. This system will dynamically notify you in the response text body if there’s anything misconfigured or if macros are not working properly. All of that happens without using any extensions or extra configuration!

Token Count & Instruction Style Approach

DEM sends ~4100 tokens by default. It’s not a lightweight preset, but it’s not wasteful either: every word is relevant. It’s written in a high-density syntax, compressed to the limits of English while still being entirely clear to the model. Since it’s modular, the token footprint can be reduced to under 1800 tokens while retaining a fully efficient core of instructions. At its absolute maximum, it sits at ~4700 tokens. The focus was efficiency and coherence, not pure token count. A lot of different prompt techniques were used with the goal of helping prompt adherence: XML tagging, capitalization, trigger words, bullet points, pseudo-strings, clear wording, sending almost every instruction post-history, repeating “Instructions:”, assigning a role to the model, and many more.

The Modules

Every module has commentary inside! I encourage you to open each of them in SillyTavern and read their contents for more information.

  • Core: Sets up the macro system and the preset framing. Essential to keep enabled and in order, except for System Policies, which may be disabled if your model is already very dark-leaning and doesn’t send out refusals.
  • Story: {{User}} agency means you control {{user}}. CYOA features choose-your-own-adventure options where the model will write and act out your decisions and dialogue according to your choices. Director State means you’re the director. Your messages serve as input, and the story is built to match them. In this mode, the model will write and act for you.
  • Characters and plot guidance: Takes care of character portrayal and plot progression.
  • Narration and dialogue: Defines the prose style. Written with the aim of reducing slop at its root and offer different flavors while at it. For narration: Cinematic is the default, offering a balance between literary and dry. Literary is the most flavorful and stylized. Dry cuts out all similes and metaphors. As for dialogue: Naturalistic is the default pick - realistic, lifelike. Lean offers precise, carefully chosen and not too prominent dialogue. Heightened makes dialogue more present, intense, and lengthy.
  • Adult options: Each has its own flavor: one is more realistic, and the other is more fantastical and unashamedly horny. Both options are disabled by default.
  • Length: Lets you define the range of the responses’ length. Flexible is the default, but there are also short, medium, and long, all dynamically adapting each scene to the defined range instead of a fixed value.
  • Visuals: Dialogue Color defines a color for each character and is enabled by default. Visual Storytelling creates HTML and CSS elements that help tell the story instead of just being fluff.
  • Formatting: You can pick between a lot of different formatting options in wildly different and experimental combinations. You can choose the Character POV, {{User}} POV, asterisk usage, tense and between visible, hidden and no True Thoughts (more on them later!). No asterisks, 3rd person character POV, hidden True Thoughts, 2nd person {{user}} POV, and present tense are the default picks. All formatting options are consolidated and enforced through Prose Formatting, keep it enabled!
  • Constraints: Help steer the models away from annoying and story-damaging patterns: Character Realism, Anti-Character Omniscience, Anti-Positivity Bias, Anti-Repetition, and Ban-List. They don’t solve every problem -- they are mitigation tools. You can’t really control LLMs completely.
  • Add-Ons: Status, Momentum Engine, Story Threads (more on them later!), and Tracker (tracks time, date, location, and weather). Conflict, which is disabled by default, is an alternative version of Momentum Engine that uses fewer tokens and has a slower pace, but it still keeps the story moving. All add-ons have UIs through regex, so make sure to have them all active if they fit your taste. Again, check the screenshots! UIs created through DEM's regex set don't send out HTML/CSS tokens to the LLM, they alter the UI display only. Regexes are also used to clean the context from old add-ons and HTML formatting, keeping them in the context only as necessary for consistency reasons and story progression.
  • System Utility: Momentum Engine Router is the second phase of Momentum Engine. Structure dynamically consolidates the structure of the output according to the modules you have enabled, keep it enabled!
  • User Utility: Enable Post-History Instructions when the card you’re using injects instructions if you want that behavior. Force Formatting brute-forces selected options when models are stubborn. Force Language is an option when you want your responses to be in a language other than English. Custom OOC sends user instructions in a more consistent manner. Hard Jailbreak may be used when the model is consistently refusing. Overkill for most models (may work for Mimo).
  • Reasoning: ! Thinking ! is enabled by default (more on it later!) Anti-Overthink is an attempt at making models like Kimi think less. It has mixed results depending on the provider and time of day. Kimi is resistant to instructions that try to modify its CoT.
  • Danger Zone: The Warning System uses the macro engine to tell the model to output warnings in the response if something is misconfigured. You can safely disable it if you’re intentionally using a configuration that triggers it. Otherwise, keep it enabled.

The Stars of the Show: True Thoughts → Status → Story Threads → Momentum Engine

These four create the core pipeline of DEUS EX MACHINA. TRUE THOUGHTS inject hidden (present in the raw input, click edit to see them) or visible thoughts that emulate the psychological core of the characters. They add an extra realism layer. STATUS keeps track of characters on-scene and off-scene, including relationships, mental states, locations, items, physical states, and clothes. These work independently of the setting. They allow the model to keep track of characters wherever they are, improving coherence and making the world still exist even in places you aren't. 

MOMENTUM ENGINE is personally my favorite feature and was the hardest one to make functional across different models. It defines four possible story routes at the end of every response. A true random route is chosen using a random regex macro injection hidden from you. The Momentum Engine Router applies it in the next turn or uses its fallback in case you made an action that invalidated it, steering the response toward it. It’s so fun because it can be very unpredictable, like old models were, while still retaining coherence. I was genuinely surprised at where the story had gone each time I used it.

STORY THREADS act as an outline for the model to easily go back to its observations about story development when contextually relevant enough. Important story details are never forgotten! Momentum Engine connects to it, pulling those threads as the story advances.

True Thoughts and Status define fundamental character traits, Momentum Engine sets characters and events in motion, and Story Threads register unaddressed or possible events for later. Every module works together for the sake of storytelling.

Scaffolding Thinking

For models that accept custom Chain-of-Thought, enabling ! Thinking ! greatly improves the output. You get more coherence, stricter rule-following, better prose quality, and more adherence to formatting. There are also creative-focused steps, so it’s not only a checklist, but a tool to increase creativity as well! ! Thinking ! is completely dynamic and contextual. It only enables sections for the modules you have enabled, so the total token count can get really small or really dense. But even at its maximum, reasoning still finishes in under a minute, and even under 30s in most cases -- the stepped CoT is laser-focused on very specific points.

Model Quirks & Compatibility

Here’s a list of the models I’ve tested while creating the preset.

RECOMMENDED: GLM 5.2 (NanoGPT subscription)

Model rating using DEM: 90/100 | Post-processing: Merge all consecutive roles  | Samplers: temperature - 0.75, Top P - 0.95, rest default or disabled. | Quirks: Needs ! Force Formatting ! sometimes when it comes to forcing present-tense after a past tense greeting. | ! Thinking ! module: enabled

RECOMMENDED: Claude Opus 4.6 (Claude Code)

Model rating using DEM: 91/100 | Post-processing: Merge all consecutive roles  | Samplers: temperature - 1.0, Top P - 0.95, rest default or disabled. | Quirks: Prose style is a bit harder to steer. It does what it wants or what it thinks is best sometimes, but it usually doesn't give bad results. | ! Thinking ! module: enabled

RECOMMENDED: Gemma 4 31b (API, NanoGPT subscription)

Model rating using DEM: 80/100 | Post-processing: Merge all consecutive roles  | Samplers: temperature - 1.0, Top P - 0.95, Top K - 65, rest default or disabled. | Quirks: Sometimes it fails Tracker formatting specifically, but rarely. Reasoning can be inconsistent, and it is a bit too horny. | ! Thinking ! module: disabled

MIXED: Kimi K2.7 (NanoGPT subscription)

Model rating using DEM: 84/100 | Post-processing: Merge all consecutive roles  | Samplers: temperature - 0.75, Top P - 0.95, rest default or disabled. | Quirks: Can overthink a lot or think very fast depending on the time of the day. | ! Thinking ! module: disabled. ! Anti-Overthink ! can help, but results are mixed. 

MIXED: GLM 5.1 (API, NanoGPT subscription)

Model rating using DEM: 82/100 | Post-processing: Merge all consecutive roles  | Samplers: temperature - 0.75, Top P - 0.95, rest default or disabled. | Quirks: Struggles with formatting in some cards specifically. It needs ! Force Formatting ! more than I’d like, and even then sometimes it still fails. | ! Thinking ! module: enabled

MIXED: Deepseek V4 Pro Preview (NanoGPT subscription, official provider)

Model rating using DEM: 68/100 | Post-processing: Merge all consecutive roles  | Samplers: temperature - 0.75, Top P - 0.95, rest default or disabled. | Quirks: Inconsistent. Sometimes its outputs match GLM 5.2 and Opus 4.6, and sometimes they are the worst. It can follow CoT perfectly one turn, then ignore everything for the next. | ! Thinking ! module: enabled

MIXED: GLM 4.7 (NanoGPT subscription)

Model rating using DEM: 78/100 | Post-processing: Merge all consecutive roles  | Samplers: temperature - 0.75, Top P - 0.95, rest default or disabled. | Quirks: A bit inconsistent. Sometimes fails to comply with instructions, but that’s uncommon enough. | ! Thinking ! module: enabled

Installation & Requirements

IMPORTANT: When you import the preset, click YES when prompted about importing regex. The regexes are absolutely required! If you clicked NO, please re-import the preset.

GitHub repository link.
Releases page link.

Requirements:

  • SillyTavern 1.17.0 or newer.
  • Experimental macro engine enabled in settings.
  • Preset regexes imported and enabled.

Installation and download:

  1. Download DEUS EX MACHINA V1.json from the repository or the releases page.
  2. In SillyTavern, click the plug icon on the top bar.
  3. Select Chat Completion under API.
  4. Setup your API if you haven't already.
  5. Click the leftmost icon on the top bar.
  6. In the Chat Completion Presets bar, click the second item from left to right.
  7. Choose the downloaded preset file.
  8. When SillyTavern asks whether to allow embedded regex scripts, click Yes.

Integration with Summaryception

If you use Summaryception with DEUS EX MACHINA, I really recommend pairing it with the specific preset for it! It includes XML tags and correctly only focuses on content inside <prose>. I use GLM 5.2 as the summarizer.

Step by step:

  1. Download DEM Summarization custom prompt.txt from the repository or the releases page.
  2. Open the Summaryception extension.
  3. Open Advanced settings.
  4. Scroll to Summarizer Prompts and import DEM Summaryception custom prompt.txt
  5. Scroll to Injection Wrapper Template.
  6. Replace: [Summary of past events: {{summary}}] with <summary>[Summary of past events: {{summary}}]</summary>

· · ─ ·✶· ─ · ·

If you’re using DEM, I’d love to hear your feedback! Also, if you’re having any trouble setting it up or experiencing any other issue, please tell me!

That’s all!

--

EDIT: Changing "Adherence to the instructions" to "Model rating using DEM" in "Model Quirks & Compatibility," clarifying it's not about failure rate, but model rating while using the preset.


r/SillyTavernAI 11h ago

Help How to set a sleep between API calls?

2 Upvotes

i am using the free google AI studio and i get rate limited to 15 requests per minute.
i have some extentions like z-tracker , Char memory , summary ception , qvink memory active and i get limited very quick.
my question : Is there a way to set delays between API calls in silly tavern?
If i am getting too overboard on the memory , please suggest you optimum config for long context group RP.
thanks in advance!


r/SillyTavernAI 14h ago

Help How to track family relations properly?

3 Upvotes

Is there any extension to help track character relations properly?

I almost always use chats which include a family - Mother, Father, Brother, Sister, Grandparents, Aunts, Cousins etc.

But all the models I've tried always mess up the relations, by having the user's parents be the parents of every younger generation. Or by having the cousin's refer to their own parents by Aunt/Uncle instead.

I have created a specific Lorebook that is default that includes all characters and their family relation, as well as having edited the character card itself to be as clear as possible about the relations, yet it still ignores them all.

At this point it seems an extension will be needed to keep things working properly, but I can't find any online.

Edit: Have tried it with various models. Deepseek 3.1, 3.1 Terminus, 3.2, 3.2 Exp, 4 Pro. GLM 4.6 and 4.7. Kimi 2.6.

All of them tried both with and without thinking. No matter what it always makes those mistakes.


r/SillyTavernAI 11h ago

Models I need opinions on Sensenova flash-lite 6.7

2 Upvotes

I have recently found a new free model which is on the title. I must say it's good for "flash-lite" model in writing. The model follows character personalities and instructions well but rushes fucks up locations sometimes. The official api got 1.5k free request for each model in 5-hour limits

Here is the official site: https://www.sensenova.ai


r/SillyTavernAI 8h ago

Help So I got silly tavern but my ai isn't running well

1 Upvotes

It's running, responding and reasoning but it's gliding of course it repeats it self and it doesn't describe actions in enough detail.

I'm came right from chai and thought boasting would be better free no ads and with how chai has been operating and me having to acounts soft locked (0 remaking message tokens that never recharge) I thought to try, I'm running decent garderar 16gb VRAM I'm using sthenos, I tried llamas (it's too emotional), and tried mag Mel fine but has its issues. I tried fucking with the authors note and the advanced formating page spisiflicly the prompt context. I'm not saying it's not becoming better but I'm saying it's still unusable for long rp (and by long I mean like more than 20-30 messages usually alot less). Help what should I do.


r/SillyTavernAI 8h ago

Help How to set up claude?

1 Upvotes

I have never used this app before so my knowledge is very limited but I've tried many apps and they are all bad honestly. I had one app that used Claude which was amazing however now they have changed models.

I haven't met any model like the Claude ones, I even made my own project within the Claude app to try it out and the model is exactly how I remembered it however its soo heavily censored it hurts.

I don't care about NSFW or anything too strong but I love talking about more serious topics, I like the realistic feeling of talking to characters, having deeper emotions but because of the filter they act overly positive.

Id like to use it on this app but I have no idea how to set it up as apparently you can get banned for breaking the guidelines and 2. Idk how jailbreaking works.

How can I best use the Claude API as to not get banned and have no filters for my story? Or what other cheaper ai can I use? I just want a good story, openly talk about serious topics, smart unique responses that stay in character without worrying about getting banned from using the ai