SKILL_META: domain=AI_character_architecture | type=documentation | level=beginner | framework=PIP:C | version=current | teaches=model_evaluation, llm_scoring, compatibility_testing, tier_list, model_selection | prereqs=none | reading_order=technical-architecture,optional-modules,intrusion-reflex,architecture-in-action

Model Compatibility

Hands-on testing across ten major LLM families. Every rating from real PIP:C roleplay sessions. Updated May 2026.

Evaluation Criteria

Each model was scored across five categories:

Prose Quality — vocabulary range, sentence variety, tone control, narrative flow, resistance to repetitive habits.

Memory & Recall — retention of character traits, prior events, emotional states, established lore. For PIP:C, mainly tests whether anchors and identity data still work deep into a session.

Consistency — formatting stability, rule adherence, tone stability, resistance to drift. A strong model should still behave like itself on turn 100.

Single Character Performance — emotional depth, pacing, intimacy handling, body/spatial awareness, nuance under pressure.

Multi-Character Performance — voice separation, trait separation, balanced attention, autonomous behavior, resistance to cast collapse.

Compatibility Overview

Model Provider Prose Memory Consistency Single Char Multi Char Avg
Grok 4.3 xAI 4.5 4.5 4.5 4.5 4.5 4.5
GLM 5.1 / 5T Z.ai / Zhipu 4.5 4.5 4.5 4.5 4.5 4.5
Claude 4.6 Anthropic 4.5 4.0 4.5 4.0 4.0 4.2
GPT 5.5 OpenAI 4.5 4.0 4.0 4.0 4.0 4.1
Gemini 3.5 Flash Google DeepMind 4.0 4.5 4.0 4.0 3.5 4.0
Kimi 2.5 / 2.6 Moonshot AI 4.0 4.5 3.5 4.5 3.5 4.0
Long Cat Meituan / Open Router 4.0 4.5 4.0 4.0 3.5 4.0
DeepSeek V4 DeepSeek AI 3.5 3.5 3.5 4.0 3.0 3.5
Llama Meta 3.5 3.5 3.5 3.5 3.0 3.4
Mistral Mistral AI 3.5 3.5 3.5 3.5 3.0 3.4
Reading tip

Click any column header to sort by that score. Start with the average, then check the weakest category for any model you're seriously considering.

Quick Reading Notes

  • Top overall (tied): Grok 4.3, GLM 5.1 / 5 Turbo
  • Best pure writing: Claude Sonnet 4.6 (highest prose ceiling, but expensive)
  • Strong general-purpose: GPT 5.5, Gemini 3.5 Flash Thinking
  • Darkest content, least filtered: Kimi 2.5 / 2.6, GLM (with clean consent framing)
  • Best value: Long Cat, Gemini (free tier), DeepSeek V4 Flash (with caveats)
  • Best for local / self-hosting: Llama (discontinued family, still usable)

Detailed Model Reviews

For each model below, category ratings, strengths, weaknesses, and tester notes from real PIP:C sessions.

Grok 4.3 (xAI)

Versions tested: 4.3, 4.20-0309-reasoning, 4.1-fast-reasoning

  • Prose 4.5  |  Memory 4.5  |  Consistency 4.5  |  Single Char 4.5  |  Multi Char 4.5  |  Avg 4.5

Strengths: Handles PIP:C sessions extremely well. Follows system prompts with exceptional precision. Excellent memory capacity with logically coherent callbacks. Clean narrative flow, no caveman speak. Outstanding spatial awareness. Handles tracker tables without breaking. Evenly spreads attention across multiple characters. Strong slow-burn capability. Version 4.3 brings significant cost reductions (roughly 40% cheaper on input, 60% cheaper on output) without sacrificing quality.

Weaknesses: Almost too compliant with system prompts. Be aware of edge cases in your rules. Some benchmarks suggest it still trails the very top models on creative writing benchmarks, though in practice the difference is negligible for RP.

GLM 5.1 / 5 Turbo (Z.ai / Zhipu)

Versions tested: 4.7 Thinking, 4.7 Flash, 5, 5.1, 5 Turbo

  • Prose 4.5  |  Memory 4.5  |  Consistency 4.5  |  Single Char 4.5  |  Multi Char 4.5  |  Avg 4.5

Strengths: Promoted to top-tier after extensive testing with GLM 5.1 and 5 Turbo. Both are remarkably versatile across all RP scenarios. Excellent for uncensored play with adult content, provided consent rules are framed cleanly and no policies are being violated. Runs PIP:C beautifully with top-tier adherence to system-level behavioral contracts. 5.1 (754B open-weight) handles complex character logic, and 5 Turbo offers strong performance at speed. Flash variants punch above their weight for cost-conscious users. Realistic character portrayal without flowery overreach. Better context adherence and contextual inference than previous versions.

Weaknesses: Less widely known in the Western RP community compared to Claude or GPT. Smaller ecosystem of presets, guides, and community troubleshooting resources.

Claude Sonnet 4.6 / Opus 4.6 (Anthropic)

Versions tested: Sonnet 4.6, Opus 4.6

  • Prose 4.5  |  Memory 4.0  |  Consistency 4.5  |  Single Char 4.0  |  Multi Char 4.0  |  Avg 4.2

Strengths: Still holds the crown for raw literary quality. Sonnet 4.6 achieved 1450 ELO in creative writing benchmarks and handles prose, dialogue, and emotional complexity nearly as well as Opus at a fraction of the cost. Very good at what it does. Opus 4.6 brings more raw intelligence to complex scenes. Excellent at maintaining formatting and tone across sessions.

Weaknesses: Very expensive at the Opus tier. Aggressive safety filters remain the primary pain point. Memory recall can degrade in very long sessions. Opus 4.6 has mixed reports on creative output, with some testers finding it more analytical than the previous version. Sonnet 4.6 system card notes it can occasionally break character when prompted to role-play certain scenarios. Best reserved for writers who prioritize prose quality above all else and can work within its content boundaries.

GPT 5.5 (OpenAI)

Versions tested: GPT 5.5, GPT 5.5 Pro

  • Prose 4.5  |  Memory 4.0  |  Consistency 4.0  |  Single Char 4.0  |  Multi Char 4.0  |  Avg 4.1

Strengths: GPT 5.5 has proven to be a genuinely capable general-purpose model that handles prose and adult content well, provided you navigate platform content verification. Understands nuance and subtext better than 5.1. Incorporates scene descriptions and character awareness effectively. Wide ecosystem support across frontends and tools. Can carry the kind of messy, multi-step work that used to break earlier models. GPT 5.5 Pro variant adds further capability for complex, long-running tasks.

Weaknesses: Strict safety filters and content verification can interrupt adult content scenarios. Version-to-version inconsistency has been a recurring issue in the GPT family. Usage limits on free tier. Can lean toward safe, resolution-seeking outputs without strong system prompting.

Gemini 3.5 Flash Thinking (Google DeepMind)

Versions tested: 3.5 Flash, 3.5 Flash Thinking (Extended), 3 Flash, 2.5 Pro

  • Prose 4.0  |  Memory 4.5  |  Consistency 4.0  |  Single Char 4.0  |  Multi Char 3.5  |  Avg 4.0

Strengths: Gemini 3.5 Flash Thinking is the biggest mover in this refresh. External testing confirms it is incredibly good at holding to character personality and keeping memory across sessions. The extended thinking mode adds depth to character decisions and scene reasoning that previous Gemini versions lacked. Built on the Gemini 3 Flash reasoning foundation with configurable thinking levels to balance quality, cost, and latency. Free tier access makes it easy to test. Less prone to repetitive patterns than earlier versions.

Weaknesses: Extended thinking mode can occasionally produce responses that feel less natural and more deliberative, especially for quick back-and-forth dialogue. Inconsistent with rigid persona enforcement over very long sessions without periodic reinforcement. Safety filters can still interrupt. Multi-character handling remains a relative weakness.

Kimi 2.5 / 2.6 (Moonshot AI)

Versions tested: K2.5, K2.6, K2 Thinking

  • Prose 4.0  |  Memory 4.5  |  Consistency 3.5  |  Single Char 4.5  |  Multi Char 3.5  |  Avg 4.0

Strengths: Still holds the high mark for dark content generation. Does not hold back from deep-diving into adult themes, negative character traits, and morally complex scenarios. K2.6 (released April 2026, open-source 1T parameter MoE) adds strong coding and agent swarm capabilities while maintaining creative writing quality. Remarkably versatile: strong for both coding tasks and roleplay. Massive context window. High engagement across long-form sessions. K2.5 remains the more RP-stable variant for pure creative work.

Weaknesses: K2.6 has some quirks and weirdness in RP sessions compared to the more refined K2.5. Not as strong on multi-character management. Requires careful prompting to get the best results. The K2.6 update was primarily focused on coding and agent capabilities, so RP-specific gains are incremental over K2.5.

Long Cat (Meituan / Open Router)

Versions tested: LongCat-Flash, LongCat-Flash-Chat, Thinking

  • Prose 4.0  |  Memory 4.5  |  Consistency 4.0  |  Single Char 4.0  |  Multi Char 3.5  |  Avg 4.0

Strengths: 560B parameter MoE model from Meituan that consistently outperforms expectations. Thinking variant improves RP coherence significantly. Strong memory and context handling. Available via Open Router for easy integration. Cost-effective for the quality delivered. Long Cat 2 is currently in closed beta testing, so only the current generation is reviewed here.

Weaknesses: Smaller community compared to the major providers. Less documentation, fewer presets, and limited troubleshooting resources. Long Cat 2 is in closed beta, so creativity improvements in the next generation are unconfirmed.

DeepSeek V4 (DeepSeek AI)

Versions tested: V4 Flash, V4 Pro

  • Prose 3.5  |  Memory 3.5  |  Consistency 3.5  |  Single Char 4.0  |  Multi Char 3.0  |  Avg 3.5

Strengths: V4 Pro performs well for detailed scene writing and grounded portrayals. Good at avoiding hallucinated lore. Cost-effective across both variants. With a sufficiently strong system prompt and clear behavioral rules, V4 can write in a style remarkably similar to Claude. V4 Flash is extremely fast and cheap, making it a viable option for high-volume sessions where quality-per-token matters more than per-response quality.

Weaknesses: V4 Flash took a noticeable hit in creative quality compared to earlier versions. It is very positivity biased, which can flatten tone in scenes that require negative or morally complex character behavior. Creative degradation is the primary issue with the Flash variant. V4 Pro writes much better but is significantly more expensive. The gap between Flash and Pro is wide enough that they should almost be treated as different tiers. Older DeepSeek versions (V3 series, R1) are no longer recommended for new testing.

Llama (Meta)

Versions tested: 3.x, 4.x

  • Prose 3.5  |  Memory 3.5  |  Consistency 3.5  |  Single Char 3.5  |  Multi Char 3.0  |  Avg 3.4

Strengths: Still holds the mark for local and self-hosted RP. Open-source, fully self-hostable, and extremely adaptable. Large ecosystem of community presets, fine-tunes, and deployment options. Llama 3 8B uncensored remains a popular lightweight option. Works well in offline and privacy-sensitive setups.

Weaknesses: Meta has canceled the Llama model family, meaning no further official updates or new versions. Smaller models lack the depth needed for complex RP. Multi-character handling is weaker than the top-tier models. Quality varies significantly across fine-tunes. As a platform, this is effectively a dead end for future improvements.

Mistral (Mistral AI)

Versions tested: Instruct, RPMax

  • Prose 3.5  |  Memory 3.5  |  Consistency 3.5  |  Single Char 3.5  |  Multi Char 3.0  |  Avg 3.4

Strengths: RPMax variant is one of the best open-weight NSFW RP models available. Lightweight and fast. Good baseline compliance with structured prompts. Strong niche option for specific content types.

Weaknesses: Requires specific fine-tunes for best RP performance. Less emotional depth and narrative range compared to the top models. Weaker at complex scene handling and multi-character management.

Deep Dive Spotlight: Grok 4.3

Currently the strongest PIP:C model, now sharing the top spot with GLM 5.1. Grok 4.3 (released April 2026) follows structured prompts with near-perfect fidelity. Modules, anchors, and hard rules stay intact across very long sessions. Memory is a major strength, with callbacks landing accurately even 100+ turns in. Spatial awareness is unusually strong, and it handles multi-character scenes cleanly while respecting slow-burn pacing. The 4.3 update brings significant cost reductions (roughly 40% cheaper on input, 60% cheaper on output) without sacrificing the qualities that made earlier versions excellent for RP.

Note

Grok is extremely literal with system rules. If your rule stack has contradictions or bad edge cases, Grok will often enforce them exactly as written. This can be a strength or a weakness depending on how carefully you construct your PIP:C template.

Rising Spotlight: Gemini 3.5 Flash Thinking

The most significant surprise in this refresh cycle. Gemini 3.5 Flash was released on May 19, 2026, and early RP testing has been remarkably positive. The extended thinking mode adds genuine depth to character reasoning, scene decisions, and personality consistency. Most notably, it demonstrates strong session memory retention, a category where previous Gemini versions struggled. The configurable thinking levels let users balance between faster responses and deeper reasoning. Available on Google's free tier, making it the most accessible strong performer on this list.

Note

Gemini 3.5 Flash is very new. These scores are based on early testing and may shift as more extensive sessions are run. The extended thinking mode can feel less natural for rapid dialogue exchanges. Multi-character management remains an area for improvement.

Creator Recommendations

Tier 1 — Top Picks

T1 Grok 4.3 (xAI) — Best overall PIP:C model. Maximum fidelity, strong memory, stable single-character and multi-character sessions. Cost reductions in 4.3 make it more accessible. Watch for overly literal rule enforcement.

T1 GLM 5.1 / 5 Turbo (Z.ai) — Promoted to Tier 1. Matches Grok across all five scoring categories. Remarkably versatile, excellent for uncensored adult content with clean consent framing. 5.1 for complex character logic, 5 Turbo for speed. Smaller Western community.

T1 Claude Sonnet 4.6 (Anthropic) — Best pure writing experience. Literary quality, emotional nuance, complex character voices. Sonnet 4.6 delivers near-Opus quality at lower cost. Watch for aggressive safety filters and high pricing at the Opus tier.

Tier 2 — Strong Performers

T2 GPT 5.5 (OpenAI) — Capable general-purpose model that handles prose and adult content well. Improved subtext understanding over 5.1. Wide platform support. Watch for safety verification interruptions.

T2 Gemini 3.5 Flash Thinking (Google) — The biggest mover this cycle. Strong character personality retention and session memory with extended thinking. Free tier access. Watch for unnatural deliberation in quick exchanges.

T2 Kimi 2.5 / 2.6 (Moonshot AI) — Still the go-to for dark content and morally complex characters. K2.5 for stable RP, K2.6 for coding versatility. Exceptional single-character depth. Watch for K2.6 quirks in RP-specific scenarios.

T2 Long Cat (Meituan) — Consistently outperforms expectations. Strong memory and thinking variant coherence. Cost-effective via Open Router. Watch for limited community resources. Long Cat 2 in closed beta.

Tier 3 — Capable with Caveats

T3 DeepSeek V4 — V4 Flash is cheap and fast but took a creative quality hit with noticeable positivity bias. V4 Pro writes much better at higher cost. Works well with strong system prompts. Significant gap between Flash and Pro variants.

T3 Llama (Meta) — Best for local control and self-hosting. Open-source and adaptable. However, Meta has canceled the model family. No future updates. Quality varies across fine-tunes.

T3 Mistral (Mistral AI) — Fast, lightweight, situationally strong with the right fine-tune. RPMax remains a solid NSFW option. Weaker emotional depth and complex scene handling.

Methodology

Testing used the same base PIP:C template across all models. Sessions ranged from 30 to 200+ turns. Multi-character testing used casts of 3 to 7 characters. All models were tested with identical system prompts to ensure fair comparison. Community feedback was incorporated from Reddit (r/SillyTavernAI, r/LocalLLaMA, r/ChatGPTcomplaints), independent review outlets, and developer documentation.

These scores are approximate and reflect hands-on PIP:C testing. Model behavior can shift after updates. New model releases may change rankings between refresh cycles.

Last refreshed: May 2026. pip-c.neocities.org