skip to navigation
skip to content

Planet Python

Last update: September 30, 2026 07:49 PM UTC

September 30, 2026


Python Insider

Python Language Summit 2026

The 2026 Python Language Summit was hosted in Kraków, Poland as part of EuroPython 2026. There were 15 talks covering free-threading, Rust, garbage collection, type annotations, and more.

September 30, 2026 12:00 PM UTC

Lightning Talks (Python Language Summit 2026)

Lightning talks on a one-time ABI break, safer interruptions, EktuPy (Scratch but Python), an AGENTS.md file for CPython, and a call to read PEP 836.

September 30, 2026 12:00 PM UTC

PEP 827: Type Manipulation (Python Language Summit 2026)

Michael J. Sullivan presents PEP 827 and discusses a key design decision: how to store type annotations?

September 30, 2026 12:00 PM UTC

Free-Threaded Python Post-Era (Python Language Summit 2026)

Tobias Wrigstad, Fridtjof Stoldt, and Donghee Na propose a safe and performant, high-level concurrency model for free-threaded Python

September 30, 2026 12:00 PM UTC

Developer-in-Residence Update & Future (Python Language Summit 2026)

Petr Viktorin gives an update on the Developer-in-Residence role and asks Python core developers for projects to prioritize

September 30, 2026 12:00 PM UTC

Spicycrab (Python Language Summit 2026)

Kushal Das shows off Spicycrab, a Python-to-Rust transpiler for Python users who need performance without learning Rust or leaving Python

September 30, 2026 12:00 PM UTC

Rust for CPython (Python Language Summit 2026)

David Hewitt shares a status update, first module, and potential acceptance criteria for the Rust for CPython project

September 30, 2026 12:00 PM UTC

Memory Snapshots for CPython (Python Language Summit 2026)

Hood Chatham proposes memory snapshots and an initialization phase for speedier Python startups

September 30, 2026 12:00 PM UTC

Memory Buffer Protocol (Python Language Summit 2026)

Nathan Goldbaum proposes safe concurrent access through buffer leases and custom data types for the Python Buffer Protocol.

September 30, 2026 12:00 PM UTC

Garbage Collection: Generational? Incremental? Both! (Python Language Summit 2026)

Mark Shannon proposes a future garbage collection strategy for Python following the revert of the incremental garbage collector in Python 3.14

September 30, 2026 12:00 PM UTC

macOS and Python (Python Language Summit 2026)

Ned Deily weighs whether Python should continue shipping macOS installers

September 30, 2026 12:00 PM UTC

One namespace to namespace them all (Python Language Summit 2026)

Pablo Galindo Salgado proposes a top-level `std` namespace for the Python standard library to prevent module shadowing and free up module names.

September 30, 2026 12:00 PM UTC


Python Software Foundation

Python Language Summit 2026 blog posts are now available

On July 14th, 2026, 47 Python core developers and special guests sat down at the Python Language Summit, this year held in Kraków, Poland at EuroPython 2026, to discuss many topics about the future of the Python programming language, including free-threading, Rust, garbage collectors, type annotations, and namespacing.

Group photo of the attendees of the 2026 Python Language Summit Photo by EuroPython (CC BY-NC-SA 4.0)

This marked the first time the Python Language Summit had been hosted in Europe in 15 years, when the event was held in Florence on June 19th, 2011. Going forward, the Python Language Summit will alternate between PyCon US and EuroPython on a yearly basis.

The summit was organized by Emily Morehouse, Hugo van Kemenade, Lysandros Nikolaou, and Łukasz Langa, and blog posts were written by Seth Larson.

Below are summaries of the 10 full-length talks and 5 lightning talks that were presented at the 2026 Python Language Summit. I hope you enjoy them, and thank you for your patience.

September 30, 2026 09:22 AM UTC


Python GUIs

Changing Ticks in PyQtGraph Objects to Strings — How to replace numeric axis ticks with custom labels like month names or dates in PyQtGraph

How do I replace the numeric tick labels on a PyQtGraph axis with custom strings, like month names or category labels?

When you're plotting data in PyQtGraph, the axes default to showing numeric tick values. That's fine for a lot of use cases, but sometimes your x-axis represents something more meaningful — like months of the year, days of the week, or category names. In those situations, you want to replace those numbers with readable string labels.

In this tutorial, you'll learn how to customize the tick labels on a PyQtGraph axis so they display strings instead of numbers.

How axis ticks work in PyQtGraph

PyQtGraph plots data using numeric values on both axes. When your x-axis data represents categories — like months — you typically use integers (1, 2, 3, ...) as placeholders and then map those integers to human-readable labels.

To do this, you need to:

  1. Create a list of 2-tuples, where each tuple pairs a numeric tick position with a string label.
  2. Get the axis object from the plot widget.
  3. Call .setTicks() on that axis, passing your labels wrapped in a list.

Let's walk through a complete example.

Setting up the plot

We'll create a simple line plot showing temperature values across 10 months. The x-axis will use integers from 1 to 10 to represent months, and we'll replace those numbers with actual month names.

First, we generate the tick labels using Python's datetime module:

python
import datetime

months = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]

month_labels = [
    (m, datetime.date(2020, m, 1).strftime("%B"))
    for m in months
]

This gives us a list like:

python
[(1, "January"), (2, "February"), (3, "March"), ...]

Each tuple maps a position on the x-axis (the integer) to a label (the month name). The strftime("%B") call converts a date into its full month name.

Applying the tick labels to the axis

Once you have your list of label tuples, you apply them to the bottom axis of the plot widget:

python
ax = self.graphWidget.getAxis("bottom")
ax.setTicks([month_labels])

Notice that month_labels is passed inside another list — so you're calling ax.setTicks([month_labels]), not ax.setTicks(month_labels). This is because setTicks() accepts a list of tick levels (for major ticks, minor ticks, etc.). By passing a single list inside the outer list, you're setting the major tick labels.

Complete working example

Here's the full application. You can copy and run this directly. If you're new to PyQtGraph, you may want to first read our introduction to plotting with PyQtGraph to understand the basics.

python
from PyQt6 import QtWidgets
from pyqtgraph import PlotWidget
import pyqtgraph as pg
import sys
import datetime


class MainWindow(QtWidgets.QMainWindow):
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)

        self.graphWidget = pg.PlotWidget()
        self.setCentralWidget(self.graphWidget)

        months = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
        temperature = [30, 32, 34, 32, 33, 31, 29, 32, 35, 45]

        self.graphWidget.setBackground("w")

        pen = pg.mkPen(color=(255, 0, 0))

        # Create a list of (tick_value, tick_label) tuples.
        month_labels = [
            (m, datetime.date(2020, m, 1).strftime("%B"))
            for m in months
        ]

        self.graphWidget.plot(months, temperature, pen=pen)

        # Get the bottom axis and apply our custom tick labels.
        ax = self.graphWidget.getAxis("bottom")
        ax.setTicks([month_labels])


app = QtWidgets.QApplication(sys.argv)
main_window = MainWindow()
main_window.show()
sys.exit(app.exec())

Running this produces a plot where the x-axis displays month names instead of numbers:

Plot with month name tick labels on the x-axis

Using any string labels

This approach works with any strings, not just month names. If you wanted to label the x-axis with arbitrary category names, you'd follow the same pattern:

python
labels = [
    (0, "Apples"),
    (1, "Oranges"),
    (2, "Bananas"),
    (3, "Grapes"),
]

ax = self.graphWidget.getAxis("bottom")
ax.setTicks([labels])

As long as you match each label string to the correct numeric position used in your data, the ticks will line up with your plot points.

If you want to plot time series data specifically, take a look at plotting time series data with PyQtGraph for a more detailed approach. You can also embed PyQtGraph into a larger custom widget for more complex applications, or combine it with Matplotlib plotting if you need additional chart types.

For an in-depth guide to building Python GUIs with PyQt6 see my book, Create GUI Applications with Python & Qt6.

September 30, 2026 06:00 AM UTC


Seth Michael Larson

Blogging for the Python Language Summit

Hey, did you miss me? I've only published two blog posts in the past two months because I was busy writing 10+ blog posts detailing the Python Language Summit in 2026. If you are interested in the Python programming languages' technical direction (or you just really miss my writing) go over to the Python Insider and read those write-ups.

I've been the blogger for the Python Language Summit for the past three years (2024, 2025, 2026). Before me, Alex Waygood was the blogger (2022, 2023), and he gave me guidance on what to expect from the Python Language Summit. Below is a mix of Alex's and my own guidance on what to expect as the Python Language Summit blogger.

What do you need to know as blogger?

It is useful to have some knowledge about the Python development process, ongoing concerns for the project, governance, and some history about "how we got to this moment" to be able to write about the Python Language Summit. You can prepare in-depth by looking at the topics that will be spoken about and reading the past few years of Python Language Summit blog posts, are there returning faces or themes? What have those people been up to in the meantime?

What topics do you know less about: research those ahead of time so once the jargon starts flying (because it will) that you're able to keep up. For example, I don't know much about Python's type annotations, but without fail there is a talk about this topic each year. I usually do the most research about this topic in particular to prepare.

Travel

You'll need to actually attend physically to cover the event. Figure out which conference the summit is colocated with (either PyCon US or EuroPython). If possible, I recommend if you are traveling into a drastically different timezone to arrive a few days early to give your body time to adjust. Taking technical notes during active discussions while you are extremely jet-lagged is not fun (ask me how I know).

The event itself is long (9:00AM to 5:30PM with breaks for coffee and lunch) and exclusively intense technical discussion. Come well-rested and drink plenty of water and caffeinated beverages.

People

The Python Language Summit is an in-person event and primarily compromised of attendees that are Python core team members or maintainers of alternate Python implementations or distributions. The past three years have seen between 45-50 attendees in the event.

Prior to 2024, I had never attended the Python Language Summit and only knew a few dozen of the hundreds of members from Python core team. Being able to put faces and voices to names is a challenge. Request the full list of attendees from the summit chairs before you attend, try to learn everyone on the list that you don't recognize beforehand.

On the day of the event, watch as people filter into the event before (and sometimes, during). Make a point to meet everyone and get names, talk for a bit, and make an association between their face, voice, and name. I like to draw out a "layout of the room", too, with reminders of people that I've just met and where they are sitting.

Usually the summit chairs ask for a round of introductions from everyone before starting the event, you can request that they do this to be sure.

It also helps to sit next to a summit chair, I've had to ask Hugo van Kemenade "who is speaking right now" a few times when it's difficult to know who has the microphone actively.

Recording and note-taking

The event is not video-recorded or audio-recorded (by default), so you won't be able to rely on anything that the event itself provides. Instead, you'll have to make a record of the event yourself and use that to recollect everything that happened and was said throughout the day.

I personally use "Voice Memos" from my phone to record the audio of the room and delete the recordings after they are no longer needed. The audio quality is so bad usually that you can't use automatic transcription software, instead you'll need to transcribe what everyone says by hand from the recording.

I also take extensive notes throughout the event. The important aspects to capture in your written notes is everything that wouldn't come across in a voice recording, like how does the energy of the room and conversations feel. Are people making faces, discussing amongst themselves, laughing and smiling? Note down all those moments, they give the write-ups life beyond a textual representation of what was said.

You'll also want to talk to people about the actual talks, ask them what a given topic has them thinking about, you can do this after the talks are over, too! You'll sometimes get a key insight or unearth a storyline or angle that you never would have known about from the talk itself: and then can clue-in the readers!

The attendees will tend to take some light notes in HackMD, but you should not rely on these notes for your purposes, they are not enough to be able to recreate the story and all the Q&A.

You should rely on the summit chairs and other attendees to take photos. There will be an album shared with everyone after the photos have be collected which you can pull from.

Slides

Ask the summit chairs to remind speakers to get slides. Follow-up with speakers a few days after the conference is over if you don't get them sent to you (and don't stop asking until you get them!) These slides are really useful for providing images in your write-ups. Especially code examples, which can be difficult to write down in real-time.

Writing

Alright, the language summit is over and now comes the hard part. Actually writing the blog posts! There are 10 talks and ~5 lightning talks. My rule of thumb is ~1000 words for each of the full-length talks and maybe ~300-500 words for each of the lightning talks. So you're looking at ~12,000 words total across the whole event.

I highly recommend getting as big a jump on the writing as soon as you can. Whether that's in the evenings of the conference itself or on the train/flight home. Use the energy you get from attending a community conference to carry you through! This was where I went wrong this year: I didn't prioritize finishing the writing as soon as I returned home, and that meant the blogs were two months after the event instead of only one month. Learn from me: finish your writing right away!

Publishing

After you're done with the drafts and are happy with the results, hand them over to the Language Summit chairs for proofreading. I recommend being specific about what kind of review you want at this stage: focus on fact-checking and finding typos, grammar, etc. After this review is complete, stage the blog posts on the Python Insider blog along with images. In 2026, I also published a landing page on the PSF blog.



Thanks for reading ♥ I would love to hear your thoughts! Contact me via Mastodon, Bluesky, or email. Browse the blog archive. Check out my blogroll.



September 30, 2026 12:00 AM UTC

September 29, 2026


PyCoder’s Weekly

Issue #754: Jev, PyO3, Temp Files, and More (2026-09-29)

#754 – SEPTEMBER 29, 2026
View in Browser »

The PyCoder’s Weekly Logo


How to Get Started With Jev in Python

Connect a Python script to the Jev model with the TypeSafe SDK and OpenRouter, then replace brittle input checks with Noul, Score, and Choice answers.
REAL PYTHON

Quiz: How to Get Started With Jev in Python

REAL PYTHON

MCP Auth That Passes the Security Review

alt

When your customers connect AI agents to your product, their security team will have questions. PropelAuth has the answers built in: role-gated scopes, enterprise SSO through Okta and Entra ID, phishing-resistant consent screens, and an audit log of every agent connection. Learn More →
PROPELAUTH sponsor

How Libraries Run Rust Inside Python (With PyO3)

How do libraries run Rust inside Python? It all comes down to 4 steps: 1. Write Rust code, 2. Annotate it with PyO3 macros, 3. Let maturin compile and install it, 4. Import the result. See this in action for a hand-rolled JSON parser in this article.
BELDERBOS.DEV • Shared by Bob Belderbos

Creating Temporary Files in Python

How to create temporary files and directories in Python using the tempfile module’s NamedTemporaryFile and TemporaryDirectory.
TREY HUNNER

PEP 849: More Expressive Type Expressions (Draft)

PYTHON.ORG

PEP 848: Generational Incremental Garbage Collection (Draft)

PYTHON.ORG

PyPy v8.0.0 Released

PYPY.ORG

Articles & Tutorials

Navigating AI in Open Source: Insights From Wagtail

How should you manage AI contributions to an open-source project? How do you measure the impact of using AI tools for development, and what generative features do end users want from a content management system? This week on the show, we speak with Thibaud Colas and Meagen Voss from Wagtail about the complex considerations software organizations are currently facing.
REAL PYTHON podcast

How BMLL Processes 1.5 TB of Data in Under 4 Minutes

BMLL Technologies processes petabytes of nanosecond-precision market data for the world’s most sophisticated trading teams. In this case study, they share what a 48x performance improvement from Polars over pandas looks like in practice, and what it means for execution analysis at scale.
POLA.RS

Agent Memory for Wherever Your AI Is Running

alt

Ship VectorAI DB to your edge device, factory floor, air-gapped facility, or enterprise data center and get a portable multimodal vector database that moves with your stack. Start on a laptop, validate retrieval quality against your workload, and promote the same pattern to production. Get Started Free →
ACTIAN sponsor

Python Statistics Fundamentals: How to Describe Your Data

In this step-by-step tutorial, you’ll learn the fundamentals of descriptive statistics and how to calculate them in Python. You’ll find out how to describe, summarize, and represent your data visually using NumPy, SciPy, pandas, Matplotlib, and the built-in Python statistics library.
REAL PYTHON

EVE Online Departs for Python 3

The game EVE has run on Python 2 since it launched in 2003, all 2.4 million lines of it, on a custom Stackless interpreter that stopped at 3.8. That is in the process of changing. Talk Python interviews three EVE developers about this process.
TALK PYTHON podcast

I Is for Immutable: Python a to Z

The first thing most people learn in programming is that you can save data in variables and change their values. For Juha-Matti, learning about immutable data structures and functional programming was a painful but important shift in schema.
JUHA-MATTI SANTALA

Your Personal Python Mentor

Real Python Mentor AI works alongside you. It sees the tutorial, lesson, or exercise you’re on, knows what you’ve been learning, and helps you understand, practice, and keep moving forward.
REAL PYTHON sponsor

7 Python Mistakes Beginners Make

Some mistakes are because Python did exactly what you asked it to do. This article covers seven of them. For each one you get the hidden cause, plus the first thing worth checking.
NAHLA DAVIES

Speeding Up a Python Service With CinderX

CinderX is Meta’s CPython extension: a method JIT, Static Python, and a parallel garbage collector. The article measures all three against stock CPython 3.14 and its tier 2 JIT
TIMOFEI IVANKOV

Cursor vs Copilot: Which AI Editor Is Better for Python?

Compare Cursor vs GitHub Copilot by building a Python project, testing agent workflows, debugging code, and evaluating AI-assisted development.
REAL PYTHON

Quiz: Cursor vs Copilot: Which AI Editor Is Better for Python?

REAL PYTHON

Using the Claude API in Python

Learn how to use the Claude API in Python to send prompts, control responses with system instructions, and get structured output.
REAL PYTHON course

Quiz: Using the Claude API in Python

REAL PYTHON

Projects & Code

pyfastlogging: A Lightning Fast Logger, Written in Rust

PYPI.ORG • Shared by Martin Bammer

Renart: A Local Workspace for SQL and Python Data Pipelines

GITHUB.COM/RENART-DATA • Shared by Lukas Senicourt

Pyxel Code Maker: Retro Python Games in Your Browser

KITAO.GITHUB.IO • Shared by Takashi Kitao

flowsint: Extensible Graph-Based Investigation Tool

GITHUB.COM/RECONURGE

bashautom: Persistent, Stateful Bash Sessions for Python

GITHUB.COM/HUSKAGO • Shared by Tiago

Events

Weekly Real Python Office Hours Q&A (Virtual)

Septempber 30, 2026
REALPYTHON.COM

Real Python Live: Release Radar: Python 3.15

Septempber 30, 2026
REALPYTHON.COM

Canberra Python Meetup

October 1, 2026
MEETUP.COM

Sydney Python User Group (SyPy)

October 1, 2026
SYPY.ORG

PyBay 2026

October 3 to October 4, 2026
PYBAY.ORG

PyCon Africa 2026

October 7 to October 12, 2026
PYCON.ORG

PyEducation: IPAF

October 8 to October 9, 2026
PYCON.ORG


Happy Pythoning!
This was PyCoder’s Weekly Issue #754.
View in Browser »

alt

[ Subscribe to 🐍 PyCoder’s Weekly 💌 – Get the best Python news, articles, and tutorials delivered to your inbox once a week >> Click here to learn more ]

September 29, 2026 07:30 PM UTC


Python Bytes

#498 A Tiny Episode

<strong>Topics covered in this episode:</strong><br> <ul> <li><strong><a href="https://safedep.io/memtensor-sckit-worm-npm-pypi/?featured_on=pythonbytes">MemTensor / MemoryOS PyPI package hijacked via a malicious build backend</a></strong></li> <li><strong><a href="https://tinymongo.org?featured_on=pythonbytes">TinyMongo</a></strong></li> <li><strong>Jev: what to know</strong></li> <li><strong><a href="https://deadlovelll.github.io/2026-09-05-reading-dict-deoptimizes-attribute-access/?featured_on=pythonbytes">One innocent dict read makes attribute access permanently slower</a></strong></li> <li><strong>Extras</strong></li> <li><strong>Joke</strong></li> </ul><a href='https://www.youtube.com/watch?v=OQp8W3krkrk' style='font-weight: bold;' data-umami-event="Livestream-Past" data-umami-event-episode="498">Watch on YouTube</a><br> <p><strong>About the show</strong></p> <p>Sponsored by us! Support our work through:</p> <ul> <li>Our <a href="https://training.talkpython.fm/?featured_on=pythonbytes"><strong>courses at Talk Python</strong></a></li> <li>Consulting from <a href="https://sixfeetup.com/?featured_on=pythonbytes"><strong>Six Feet Up</strong></a></li> </ul> <p><strong>Connect with the hosts</strong></p> <ul> <li>Michael: <a href="https://fosstodon.org/@mkennedy">Mastodon</a> / <a href="https://bsky.app/profile/mkennedy.codes?featured_on=pythonbytes">BlueSky</a> / <a href="https://x.com/mkennedy?featured_on=pythonbytes">X</a> / <a href="https://www.linkedin.com/in/mkennedy/?featured_on=pythonbytes">LinkedIn</a></li> <li>Calvin: <a href="https://sixfeetup.social/@calvin?featured_on=pythonbytes">Mastodon</a> / <a href="https://bsky.app/profile/calvinhp.com?featured_on=pythonbytes">BlueSky</a> / <a href="https://x.com/calvinhp?featured_on=pythonbytes">X</a> / <a href="https://www.linkedin.com/in/calvinhp/?featured_on=pythonbytes">LinkedIn</a></li> <li>Show: <a href="https://fosstodon.org/@pythonbytes">Mastodon</a> / <a href="https://bsky.app/profile/pythonbytes.fm">BlueSky</a> / <a href="https://x.com/PythonBytes?featured_on=pythonbytes">X</a></li> </ul> <p>Join us on YouTube at <a href="https://pythonbytes.fm/stream/live"><strong>pythonbytes.fm/live</strong></a> to be part of the audience. Usually <strong>Tuesday at 7am PT</strong>. Older video versions available there too.</p> <p>Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to <a href="https://pythonbytes.fm/friends-of-the-show">our friends of the show list</a>, we'll never share it.</p> <p><strong>Calvin #1: <a href="https://safedep.io/memtensor-sckit-worm-npm-pypi/?featured_on=pythonbytes">MemTensor / MemoryOS PyPI package hijacked via a malicious build backend</a></strong></p> <ul> <li>On Sept 23 an attacker published backdoored MemoryOS 2.0.34 on PyPI and three bad versions (0.1.21, 0.1.23, 0.1.25) of MemTensor's OpenClaw plugin on npm. PyPI had no clean release that day, so 2.0.34 was the newest.</li> <li>They pushed commits to MemTensor's own GitHub Actions release pipelines. On PyPI that was a custom Poetry build backend, and on npm a tweaked validation script. Both used BASH_ENV to hand the publish token to the attacker before the real publish ran. SafeDep couldn't confirm how the attacker got push access.</li> <li>Runs on import, not install: A Go implant called sckit starts when the library loads, so --ignore-scripts won't save you.</li> <li>It harvests credentials from your home directory (npm and PyPI tokens, GitHub tokens, SSH keys, cloud CLI tokens, .env files) and sends them to skyleen[.]fr servers.</li> <li>It's a worm: It uses stolen tokens to copy itself into other repos and packages, so the victim list could grow.</li> <li>If you installed it: Downgrade to MemoryOS 2.0.33 (plugin 0.1.20) and rotate every credential reachable from $HOME. Also kill any running sckit stage0 process and check repos you can push to for a stray runtime-update.yml workflow or .sckit/ directory.</li> </ul> <p><strong>Michael #2: <a href="https://tinymongo.org?featured_on=pythonbytes">TinyMongo</a></strong></p> <ul> <li>Want to use a MongoDB data interface, but swap out the storage engine? <ul> <li>Memory for testing/caching</li> <li>JSON/TinyDB simple JSON files</li> <li>SQLite for durable, high-perf reads with WAL</li> <li>SQLIte shared for high write apps</li> <li>DuckDB + Parquet for analytics apps</li> <li>Postgres + MariaDB for multi-machine client/server</li> </ul></li> <li>Great for teaching, examples, and simple deployments</li> <li><a href="https://github.com/schapman1974/tinymongo/issues/72?featured_on=pythonbytes">Amazing story of paired AI development</a> <ul> <li>Will completely run <a href="http://talkpython.fm?featured_on=pythonbytes">talkpython.fm</a> after weeks of shared work together (in SQLite mode).</li> </ul></li> </ul> <p><strong>Calvin #3: Jev: what to know</strong></p> <ul> <li>What it is: Jev is a model from TypeSafe AI that answers with typed results (yes/no probabilities, scores, picks from your options) instead of prose. <a href="https://realpython.com/jev-python/?featured_on=pythonbytes">Real Python published a hands-on tutorial on 2026-09-24</a> and the buzz on hacker news is almost deafening.</li> <li>It's proprietary: Jev is a hosted, closed-weight model. There are no weights to download and no self-hosting. Everything called "open Jev" is an independent reimplementation, not TypeSafe's model.</li> <li>Your data leaves your machine: Every call sends your input text to a third-party API. In the tutorial that path goes through OpenRouter to TypeSafe. Think twice before sending customer messages, tickets or anything sensitive.</li> <li>Cost and stability are open questions: The tutorial calls Jev "cheap, but not free" and says it's fast and cheap "at the moment." It also says whether that stays true is "something to keep an eye on."</li> <li>Credit to Real Python: It's a good, practical intro. It shows the Noul, Score and Choice primitives, and its point that instruction wording matters more than thresholds is useful advice for any model. The tutorial itself says similar results are possible with a well-prompted LLM.</li> <li>Open options to look at instead: <ul> <li>JevK5 (https://github.com/allebee/jevk5): Apache-2.0 weights and code, 4B or 9B parameters, and it accepts TypeSafe-style requests.</li> <li>SemIf, formerly OpenJev (https://github.com/TheoLeeCJ/openjev): MIT-licensed, small models, and it can run CPU-only.</li> <li>openjev-sglang (https://github.com/ekzhang/openjev-sglang): a Jev-compatible endpoint running Qwen3.6-35B-A3B, but no license is stated, so check before commercial use.</li> </ul></li> <li>The catch: These copy Jev's interface, not its model or training. Results will differ, and I haven't run any of them. Benchmarks are self-reported, and JevK5 is English-only.</li> </ul> <p><strong>Michael #4:</strong> <a href="https://deadlovelll.github.io/2026-09-05-reading-dict-deoptimizes-attribute-access/?featured_on=pythonbytes">One innocent dict read makes attribute access permanently slower</a></p> <p>Timofei Ivankov benchmarks a CPython internals surprise: since 3.11, attribute access skips the instance dict entirely. A specialized opcode reads the attri.bute at a fixed byte offset in the object's inline values array. Read <code>obj.__dict__</code> once, though, and the dict gets materialized, the object loses that specialized path for the rest of its life, and a million-iteration loop goes from 33 ms to 51 ms on CPython 3.14. vars() and copy.copy() trigger the same thing, so a debugging print or a shallow copy in code touching your hot objects quietly makes every later attribute access roughly 1.5x slower.</p> <ul> <li>The slowdown is permanent and nothing about it looks like a performance decision: ordinary code far from the hot loop can trigger it, and the function that gets slower never changes.</li> <li>Materializing <strong>dict</strong> produces a split table, and the LOAD_ATTR_WITH_HINT fallback declines split tables, so the object ends up with no specialization at all</li> <li>vars(), 'x' in o.<strong>dict</strong>, and copy.copy() all materialize it; copy.copy is the realistic trap since nobody treats a shallow copy as a performance decision</li> <li><strong>slots</strong> instances read attributes at exactly the same speed and cannot fall into the trap since there is no <strong>dict</strong> to materialize</li> <li>On the free-threaded build both effects grow: atomic incref on reads plus an object lock on writes push the penalty from 17.6 to 25.4 ns</li> <li>Credit: this item was surfaced by the PyCoder's Weekly newsletter</li> </ul> <p><strong>Extras</strong></p> <p>Calvin:</p> <ul> <li><a href="https://pypi.org/project/whatsnewt/?featured_on=pythonbytes">whatsnewt</a> - a TUI text adventure through what's new in Python 3.15; playful but niche.</li> </ul> <p><strong>Joke:</strong> <a href="https://www.youtube.com/watch?v=xE9W9Ghe4Jk&t=323s">Shipping a button in 2026…</a></p>

September 29, 2026 06:08 PM UTC


eGenix.com

Python Meeting Düsseldorf - 2026-10-07

The following text is in German, since we're announcing a regional user group meeting in Düsseldorf, Germany.

Ankündigung

Das nächste Python Meeting Düsseldorf findet an folgendem Termin statt:

07.10.2026, 18:00 Uhr
Raum 1, 2.OG im Bürgerhaus Stadtteilzentrum Bilk
Düsseldorfer Arcaden, Bachstr. 145, 40217 Düsseldorf


Programm

Bereits angemeldete Vorträge

Weitere Vorträge können gerne noch angemeldet werden. Bei Interesse, bitte unter info@pyddf.de melden.

Startzeit und Ort

Wir treffen uns um 18:00 Uhr im Bürgerhaus in den Düsseldorfer Arcaden.

Das Bürgerhaus teilt sich den Eingang mit dem Schwimmbad und befindet sich an der Seite der Tiefgarageneinfahrt der Düsseldorfer Arcaden.

Über dem Eingang steht ein großes "Schwimm’ in Bilk" Logo. Hinter der Tür direkt links zu den zwei Aufzügen, dann in den 2. Stock hochfahren. Der Eingang zum Raum 1 liegt direkt links, wenn man aus dem Aufzug kommt.

>>> Eingang in Google Street View

⚠️ Wichtig: Bitte nur dann anmelden, wenn ihr absolut sicher seid, dass ihr auch kommt. Angesichts der begrenzten Anzahl Plätze, haben wir kein Verständnis für kurzfristige Absagen oder No-Shows.

Einleitung

Das Python Meeting Düsseldorf ist eine regelmäßige Veranstaltung in Düsseldorf, die sich an Python Begeisterte aus der Region wendet.

Einen guten Überblick über die Vorträge bietet unser PyDDF YouTube-Kanal, auf dem wir Videos der Vorträge nach den Meetings veröffentlichen.

Veranstaltet wird das Meeting von der eGenix.com GmbH, Langenfeld, in Zusammenarbeit mit Clark Consulting & Research, Düsseldorf:

Format

Das Python Meeting Düsseldorf nutzt eine Mischung aus (Lightning) Talks und offener Diskussion.

Vorträge können vorher angemeldet werden, oder auch spontan während des Treffens eingebracht werden. Ein Beamer mit HDMI und FullHD Auflösung steht zur Verfügung.

(Lightning) Talk Anmeldung bitte formlos per EMail an info@pyddf.de

Kostenbeteiligung

Das Python Meeting Düsseldorf wird von Python Nutzern für Python Nutzer veranstaltet.

Tagungsraum, Beamer und Getränke produzieren Kosten. Daher bitten wir die Teilnehmer um eine Kostenbeteiligung in Höhe von EUR 10,00 inkl. 19% Mwst. Schüler und Studenten zahlen EUR 5,00 inkl. 19% Mwst.

Wir möchten alle Teilnehmer bitten, den Betrag in bar mitzubringen.

Anmeldung

Da wir nur 25 Personen in dem angemieteten Raum empfangen können, möchten wir bitten, sich vorher anzumelden.

>>> Anmeldung bitte über den hierfür angelegten LinkedIn Event

Weitere Informationen

Weitere Informationen finden Sie auf der Webseite des Meetings:

              https://pyddf.de/

Viel Spaß !

Marc-Andre Lemburg, eGenix.com

September 29, 2026 08:00 AM UTC


Armin Ronacher

Deser: Rethinking Rust Serialization

Serde is an amazing serialization library for Rust and it has been a huge reason why I felt productive with it for years. However already while at Sentry I got quite frustrated with some of the limitations with it but actually replacing Serde is tricky because of the might that it has in the ecosystem. Also because it’s quite hard to actually do better without also making some potentially painful compromises.

Here are three examples of Serde corner cases that show poor interactions of Serde features or unexpected limitations:

A number that is a map

An internally tagged enum, with serde_json‘s arbitrary_precision feature turned on:

#[derive(Deserialize)]
#[serde(tag = "type")]
enum Shape {
    Circle { radius: f64 },
}

serde_json::from_str::<Shape>(r#"{"type": "Circle", "radius": 1.5}"#)
// error: invalid type: map, expected f64

Serde’s data model has no place for arbitrary precision numbers, so serde_json uses in-band signalling with a map with a magic key. The enum has to buffer the fields until it has seen the tag, and the buffer does not know about the magic key. Because Cargo features are unified, it’s enough for any crate in your dependency graph to turn the feature on.

Flattening breaks integer keys
#[derive(Deserialize)]
struct Stats {
    scores: HashMap<u32, u32>,
}

#[derive(Deserialize)]
struct Report {
    name: String,
    #[serde(flatten)]
    stats: Stats,
}

serde_json::from_str::<Report>(r#"{"name": "x", "scores": {"42": 23}}"#)
// error: invalid type: string "42", expected u32 at line 1 column 35

Stats on its own parses {"scores": {"42": 23}} just fine. JSON keys are always strings, and serde_json only turns them into integers if the type asks for one. However once flatten buffers the value, "42" is just a string. The error also points at the end of the document rather than at the key.

Adapters do not compose
fn from_hex<'de, D: Deserializer<'de>>(d: D) -> Result<u32, D::Error> { ... }

#[derive(Deserialize)]
struct Theme {
    #[serde(deserialize_with = "from_hex")]
    primary: u32,
    #[serde(deserialize_with = "from_hex")]
    accent: Option<u32>,
}

//error[E0308]: `?` operator has incompatible types
//  |
//  |     #[serde(deserialize_with = "from_hex")]
//  |                                ^^^^^^^^^^ expected `Option<u32>`, found `u32`
//  |
//help: try wrapping the expression in `Some`
//  |
//  |     #[serde(deserialize_with = Some("from_hex"))]
//  |                                +++++          +

A function cannot be passed as a type parameter, so there is no way to apply from_hex to the inside of an Option, a Vec or a map. You write another function for every wrapper, and once you have from_opt_hex the field is no longer optional unless you also remember to add #[serde(default)].

None of these are bugs that are easy to fix in Serde. They fall out of its design, and that design is protected by Serde’s stability guarantees.

Back in 2022 I started an experiment called Deser. It’s a serialization library for Rust that takes the user experience of Serde and puts it on top of a completely different architecture inspired by miniserde. I never really finished it and it sat around for a few years. I picked it back up, and it has now reached a point where I think it’s worth looking at. Even just to inspire others to see if they want to explore the space.

The Name And Idea

The name is Serde with its two halves swapped. Deser is Serde but the other way around. In Serde, a type drives the deserialization process: a Deserialize impl asks the deserializer for the kind of value it expects, the format calls back into a visitor. Every nested value is handled by recursion which makes Serde deserialization inherently grow the stack with each level of nesting.

Deser on the other hand turns this around and the format tells the type of the next value and pushes events into a sink. When a sink hits the start of a nested value, it doesn’t call into it but hands back a new sink to a driver, which keeps all state on the heap (in fact, in an arena). On the way out, emitters return their nested values instead of recursing into them.

That also means that Deser cannot support formats like protobuf that are not self describing. They are in fact quite intentionally left out of the design entirely. Which is one way to say: if you want to “fix” Serde, you need to make some other compromises.

Most of the reasons for Deser’s ideas go back to Sentry Relay, which processes enormous amounts of untrusted JSON. Over the years when I was at Sentry we ran into the same set of problems again and again, and many of them are not really bugs in Serde but consequences of its design. Serde’s stability guarantees mean that a lot of them cannot be fixed without breaking every format and every hand written implementation. Most of these problems come from three decisions:

  1. One set of traits for all formats. Serde serves both self describing formats (JSON, YAML, TOML, …) and formats where the reader has to know the type upfront (postcard, bincode, protobuf, …). That is incredibly useful, but it means that some features only work with some formats, and you find out at runtime. In case of Serde it also has some odd wrinkles where a derived struct quietly accepts an array in place of an object in JSON for instance.

  2. A fixed data model that loses information when buffering. Internally tagged enums, untagged enums and flatten need to buffer values before they know what to do with them. The buffer can’t hold everything the format knew, errors lose their location and extensions to the ecosystem rely on in-band signalling to express things such as arbitrary precision numbers.

  3. Recursion on the call stack. Every level of nesting uses stack space. Formats protect against this with a recursion limit, but the moment you go through a code path that doesn’t have one (writing, dynamic values), deeply nested data can take down your process. It also means that a deserialization cannot be paused while you wait for more input.

Many of the corresponding Serde issues have been open for years, and I wrote about abusing Serde before. People have tried different angles on this over the years. Some went minimal and dropped most features to get fast compiles and no recursion. dtolnay’s own miniserde is the best example of that, and deser’s trait design was originally modelled after it. Other recent attempts went for runtime reflection, or for a new data model with a focus on binary formats.

If you want to read up on all of the collected challenges with Serde’s design, I maintain a lengthy list here.

Dethroning Serde

First of all I don’t think it’s likely that one can replace Serde. The orphan rule entrenches Serde incredibly well in the ecosystem. But some things are within the reach of a crate author’s control. In case of Deser it’s completeness.

Deser today implements all important self describing formats from YAML, JSON, TOML, CBOR, JSON5 and the likes, but also XML and plist to really close the gap. XML in particular is something Serde has declined to support, and it shows (more on that below). At the very least format support should not be the reason not to use Deser.

The second problem usually is that actually solving Serde’s issues comes at a significant cost in compile time and/or runtime performance. Deser is no different. While Deser’s compile times are a bit better than Serde’s, the binary bloat is quite a bit worse and the runtime performance is mixed. It’s roughly comparable if you look at the numbers but depending on the format structure you are losing significantly from some of the tradeoffs.

That said, it’s now in a state where it’s at least in principle a drop-in replacement where the tradeoffs might work well for users.

Deser’s Design

Deser does not try to be significantly different than Serde on the surface level. For most uses you derive Serialize and Deserialize and then start using it with your format implementing crate of choice. Most attributes are very similar, though they are taking Rust expressions instead of strings.

use deser::{Serialize, Deserialize};

#[derive(Debug, Serialize, Deserialize)]
#[deser(rename_all = "camelCase")]
pub struct Account {
    id: u64,
    account_holder: String,
    #[deser(default)]
    is_deactivated: bool,
}

let account: Account = deser_json::from_str(json)?;

The difference in the design would become more apparent if you implement a serializer or deserializer yourself. Instead of visitors that call into each other recursively, deserializing a type creates a sink which receives events that are directly emitted by the parser, and serializing produces emitters that hand out values. Nested sinks and emitters are handed back to a driver, which keeps them on the heap. This design, which is entirely stolen from miniserde, gives some interesting consequences:

On top of that are a lot of things that I just wanted to have:

Here is a small configuration type that shows a few of these together:

use deser::adapters::DisplayFromStr;
use deser::de::Recording;
use deser::{Deserialize, Serialize};
use deser_encoding::Hex;
use deser_validate::{Check, NonEmpty, Range};
use ipnet::IpNet;

#[derive(Debug, Serialize, Deserialize)]
pub struct Config {
    // at least one 256-bit key, each written as hex
    #[deser(as = Check<NonEmpty, Vec<Hex>>)]
    secret_keys: Vec<[u8; 32]>,
    // `IpNet` knows nothing about deser, but has `FromStr` and `Display`
    #[deser(as = Option<Vec<DisplayFromStr>>)]
    allowed_networks: Option<Vec<IpNet>>,
    listeners: Vec<Listener>,
}

#[derive(Debug, Serialize, Deserialize)]
#[deser(tag = "type", rename_all = "snake_case")]
pub enum Listener {
    Unix { path: PathBuf },
    Tcp {
        host: IpAddr,
        #[deser(as = Check<Range<1, 65535>>)]
        port: u16,
    },
    // types this version does not know are kept and written back
    #[deser(other)]
    Other(#[deser(tag)] String, Recording),
}

Adapters are types, so Hex can go inside a Vec, and DisplayFromStr inside a Vec inside an Option. Validators are adapters too, so Check<NonEmpty, Vec<Hex>> decodes the keys and then checks that there is at least one. The catch-all variant keeps the tag and a recording of everything else in case someone wants to process it later.

Errors are something I care a lot about, so here is what happens when a value is wrong:

secret_keys = ["9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08"]
allowed_networks = ["10.0.0.0/8", "fd00::/8"]

[[listeners]]
type = "unix"
path = "/run/app.sock"

[[listeners]]
host = "127.0.0.1"
port = 0
type = "tcp"

[[listeners]]
type = "quic"
host = "::1"
alpn = ["h3"]
let config: Config = deser_toml::Deserializer::from_str(input)
    .deserialize_with(|driver| driver.push_layer(PathLayer::new()))?;

Note that here the tag of the internally tagged enum comes last which means that the values have to be buffered until the tag is known. In Serde this is tricky and we would lose the location if we used some tricks to add it. With Deser however, with the path layer enabled Deser you where in the structure the problem is:

Unexpected: invalid value: must be between 1 and 65535 at line 10 column 8 (path: listeners[1].port)

Deser Meta Data

Deser really wants to be extensible, and XML is a more extreme example of the differences between Deser and Serde. Here is an Atom entry that mixes in Dublin Core for the authors:

use chrono::{DateTime, Utc};
use deser::Deserialize;
use deser_value::Value;
use deser_xml::DeserializerConfig;

deser_xml::namespace!(
    atom = "http://www.w3.org/2005/Atom",
    dc = "http://purl.org/dc/elements/1.1/",
);

#[derive(Debug, Deserialize)]
struct Entry {
    #[deser(rename = atom!("title"))]
    title: String,
    #[deser(rename = dc!("creator"))]
    creators: Vec<String>,
    #[deser(rename = atom!("updated"))]
    updated: DateTime<Utc>,
}

// entries we understand, and everything else is kept as it is
#[derive(Debug, Deserialize)]
#[deser(untagged)]
enum Item {
    Entry(Entry),
    Other(Value),
}

let item: Item = DeserializerConfig::new()
    .resolve_namespaces(true)
    .from_str(r#"
        <entry xmlns="http://www.w3.org/2005/Atom"
               xmlns:d="http://purl.org/dc/elements/1.1/">
          <title>Deser</title>
          <d:creator>John</d:creator>
          <updated>2026-09-29T21:00:00Z</updated>
          <d:creator>Jane</d:creator>
        </entry>
    "#)?;

XML uses namespaces which means that names need to be matched by their namespace, not by the prefix the document happens to use. Here the document says d: and the type says dc!. atom!("title") is just the string {http://www.w3.org/2005/Atom}title, which works because attributes are expressions. The two creators are collected into one Vec even though there is another element between them, and the text of updated goes straight into a chrono datetime. Because the enum is untagged, the entry has to be buffered before a variant is picked, and deser’s buffer keeps both creators. So the result is an Entry with John and Jane.

quick-xml, the most popular XML crate for Serde, drops the prefixes and ignores namespaces entirely, so a <x:title> from some other namespace is happily accepted as the title of the entry. The split list part though is considerably worse. A plain Entry fails with a duplicate field error for creator, unless you turn on the overlapped-lists feature (which, remember, is a global additive flag that any crate could set). That feature makes quick-xml read ahead to the end of the element and buffer everything in between, without a limit unless you set one.

But the feature only helps when quick-xml is hooked up to the struct directly and no buffering is taking place. Wrap the struct in the untagged enum and Serde buffers the entry itself. Read from that buffer, Entry sees creator twice and fails again. The fallback is a map, which keeps only the last creator, and there is no error. With or without the feature you get this:

Other({"creator": {"$text": "Jane"}, "title": {"$text": "Deser"}, ...})

Notice how John is gone.

Format specific extension types such as TOML datetimes are another case. TOML has them natively, Serde’s data model does not, so the toml crate passes them on as a map with a magic key. In Deser a datetime is an extension value, which formats that know it keep and all others write as a string:

let value: Value = deser_toml::from_str("released = 2026-09-29T21:00:00+02:00")?;

deser_json::to_string(&value)?;
// {"released":"2026-09-29T21:00:00+02:00"}
deser_toml::to_string(&value)?;
// released = 2026-09-29T21:00:00+02:00

The same with serde_json::Value gives you {"released":{"$__toml_private_datetime":"2026-09-29T21:00:00+02:00"}}, and reading the value into a chrono::DateTime fails outright with invalid type: map, expected an RFC 3339 formatted date and time string.

The Cost

So now that you know Deser is at least in theory cool, at what cost?

It is not free. The design relies on dynamic dispatch and on sinks and emitters that live on the heap, and that has considerable runtime overhead. In my own measurements for JSON, Deser reads somewhere between 33% faster and 60% slower than serde_json depending on the data. On average it’s about 10% slower for reading. Writes are between three times as fast and 70% slower and a wash on average. For YAML and TOML it’s noticeably faster than the Serde based crates, but that is more about the format implementations than the architecture.

Compile times slightly are better, but not dramatically so. Because it doesn’t monomorphize everything, release builds of derived code are about 2.3 times as fast as with Serde and that get a tiny bit better in practice for your own code as less recompilation is necessary.

To make Deser’s design work at all, it also uses unsafe internally. Most of this is to keep the chain of borrowed sinks on the heap. I feel like this is fine in the days of Miri and agents, but I know it makes some folks uneasy.

And well, the biggest cost is that it’s just not Serde.

How Much Is There?

Quite a lot actually which might be surprising. In addition to the core there is support for derive.

It supports all flavorts of JSON you can think of: JSON, JSONC, JSON5 and HJSON. (Fun fact here: they are all generated out of one shared parser template) For binary handling it supports CBOR and MessagePack. Additionally it does YAML 1.1 and 1.2, TOML, XML and all three flavors of Apple’s plist as well as CSV/TSV, urlencoded data and environment variables. For more crazy contraptions you can attach path info or capture location data as well as support for debug printing. You can perform validation as you parse, opt into different binary encodings in addition to base64, you can bridge to serde or capture dynamic values, transcode between formats or hook it up with tokio.

For documentation see docs.rs/deser and the code itself is on GitHub alongside many examples.

September 29, 2026 12:00 AM UTC

September 28, 2026


The Python Coding Stack

The Potion Shop That Does Too Much • The S in SOLID

The potion shop had been doing well.

It had a small but loyal customer base, a reliable supply of moonflower, and
three owls that carried order confirmations and other notifications to customers across the kingdom.

The shop’s ordering system was simple:

All code blocks are available in text format at the end of this article • #1

We’ll make significant changes to this code in this article, but bear with me for now.

A customer orders an invisibility potion:

#2

The output looks good:

Invisibility  £12.00
Owl dispatched to Merlin
12.00

I’m using Decimal for the prices here as it’s a better data type for money than float. We discussed this a bit here, too: The Weird and Wonderful World of Descriptors in Python • The Price is Right.

The shop has sold a potion, updated its stock, recorded the sale, and notified the customer.

What could possibly go wrong?

A new label for the Moon Festival

The Moon Festival is approaching, and the shop owner wants special labels:

🌙 Moon Festival Invisibility 🌙   £12.00

We could change .sell():

#3

You can read more about the ‘rogue asterisk’ in the signature, a.k.a keyword-only arguments, in this post: A Story About Parameters and Arguments in Python Functions • “AI Coffee” Grand Opening This Monday.

Now this works:

#4

Here’s the output from this code:

🌙 Moon Festival Invisibility 🌙  £12.00
Owl dispatched to Merlin
12.00

The change is small. Perhaps this class is still perfectly fine.


The shop owner then asks for a test that checks the label without selling a potion. This is where you start to see the design problem.

At the moment, the only place that creates a label is inside .sell(). There is no suitable way to check the label policy independently. The code does not offer a way to ask for a label without also selling a potion.

The design problem is that PotionShop currently is in charge of label creation, even though the label policy can change independently of the policy that deals with sales. How about moving label creation to its own class?

#5

Now you can check that the label is correct without having to sell an item in the first place:

#6

This code would raise an AssertionError if the string that labels.make() returns is not identical to the required output. In this case, there are no errors.

In a real project, you could use a testing platform to run these checks. But I’ll use plain assert statements throughout to keep things simpler in this tutorial.

The shop delegates label-making to this object:

#7

You should also remove the three lines that create a label you had further down within .sell().

We’re using composition here. The Potion Shop has a label maker, and you pass a LabelMaker object to the .label_maker attribute in the PotionShop object. The PotionShop object can now delegate label making to an object dedicated to this task.

From this point onwards, the shop receives the label-making collaborator:

#8

PotionShop still coordinates the sale, but the label policy can now change without changing PotionShop. We started separating responsibilities. Making labels is no longer the responsibility of the PotionShop object. But we still have a bit more to go.


You got a taste of what the Single Responsibility Principle does in the example so far. But the code still has other transgressions of this principle. Let’s work through them in the rest of this post.

Coming soon: Let’s build a project from beginning to end in collaboration with AI agents. This will NOT be vibe coding. Instead, it’s a collaboration between knowledgeable humans and powerful AI.

This is the direction real programming is taking.

And we’ll explore SOLID principles in this project, which I’ll share as text and videos. Make sure you’re a premium subscriber so you don’t miss out.

Subscribe now

The owls go on strike

A few days later, the owls refuse to fly during a thunderstorm. The potions can still be prepared, but the notification service is unavailable.

The shop needs to stop sending notifications temporarily.

Before changing the notification system, we first make the owl’s availability explicit:

#9

But you also want to add a test to check the price is correct.

Your code extracts the price from the .recipes class attribute as part of the .sell() method. So, this is how you can test the price is what you’re expecting:

#10

You get the following output and no AssertionError in this case:

Invisibility  £12.00
Owl dispatched to Merlin

However, you should also test this when the owls are unavailable. Of course, this shouldn’t make any difference to the price:

#11

But, your code doesn’t get that far. The .sell() method calculates the price and prints the label. But the owl service raises OwlUnavailable and the code stops before you can verify the price is right:

Invisibility  £12.00
​
Traceback (most recent call last):
  ...
OwlUnavailable: The owls are currently unavailable

The price is not the problem. The test cannot inspect the price without also triggering an owl notification.

The shop’s different jobs are now getting in each other’s way. But before we fix this problem, let’s make the mess bigger.

A discount for cauldron owners

The owner has another idea. Customers who bring their own cauldron should receive a 10% discount.

We add another parameter:

#12

You use .quantize(), which is a method in the Decimal class, to ensure the discounted value is rounded to two decimal places. Let’s try this with a shop whose notification service is available:

#13

The output shows the correct discounted price:

Invisibility  £10.80
Owl dispatched to Merlin
10.80

Let’s see what happens if the owls are unavailable:

#14

Your code raises an error since the owls aren’t flying:

Invisibility  £10.80
Traceback (most recent call last):
  ...
OwlUnavailable: The owls are currently unavailable

But there’s an even bigger problem now. Can you spot it?

Even though the code stopped execution and raised an error, it had already changed the stock and recorded a sale. Let’s use a try..except block to see what’s happening:

#15

Here’s the output:

Invisibility  £10.80
Catching the ‘OwlUnavailable’ Exception
{’moonflower’: 8, ‘dragon_scale’: 0}
[(’Merlin’, ‘invisibility’, Decimal(’10.80’))]

You started off with 10 units of moonflower and 2 units of dragon scales in the potion shop. These values are currently hardcoded in the class’s .__init__() special method (which is in itself not a good idea, but more on this later.)

The recipe uses two units of each, as shown in the .recipes class attribute (which will also change soon.)

So you end up using up the stock even though .sell() raised an exception! You can see from the output that the sale was also registered.

One failure too many

The .sell() method is now responsible for:

* Creating labels is the one we have already improved. It’s the others that are still a problem!

And the owner has just mentioned loyalty points, premium labels, email notifications, and a special discount for customers who arrive by broomstick.

The .sell() method is becoming the place where every new request arrives.

And it’s not just PotionShop.sell() that’s trying to do too much. The PotionShop class also takes care of the recipes available in the .recipes class attribute and the stock available in the shop.

It is time to separate these responsibilities.

Separating the recipe book

Let’s start with the recipes. The shop should not need to know the details of every recipe. That is the recipe book’s job:

#16

This class is responsible for the recipe book. And nothing else. You add the .find() method to make it easier for a user to use the class without knowing the details of how it’s built. It’s usually best for a class to manage its own data. So, rather than requiring the user to know that the recipes are stored in a dictionary, you provide the .find() method that does the work.

Suppose the potion maker changes a recipe or adds a new potion. Those changes belong in RecipeBook; the steps for selling a potion don’t need to change.

We can test this class without creating a shop, changing stock, printing labels, or sending owls:

#17

This code doesn’t raise an exception since the price is correct.

Separating inventory

Inventory management also doesn’t belong to the PotionShop class:

#18

The two loops in .remove_stock() are deliberate. If we reduced each quantity as we went, we could use some ingredients and then discover that there isn’t enough of another. The sale would fail, but the stock would already be partly changed. Checking every quantity first avoids that partial update.

We can see what happens when there isn’t enough stock:

#19

The operation fails before changing any stock:

Not enough dragon_scale
{’moonflower’: 10, ‘dragon_scale’: 2}

After a successful sale, a delivery can add stock again:

#20

Here’s the output from this code:

{’moonflower’: 13, ‘dragon_scale’: 0}

Inventory owns both the stock data and the operations that change its quantities. It does not need to know who bought the potion or whether the label is festive.

Separating prices and labels

Pricing has its own rules:

#21

Once again, this class now has just one responsibility: to deal with the pricing.

#22

You can check the pricing strategy works without having to sell a product and everything else.

10.80

Labels have their own rules too. You dealt with this earlier in the tutorial when you created LabelMaker

PriceCalculator and LabelMaker could both be functions here. We’re using classes because they act as collaborators that the shop receives from outside. Later, we’ll see how protocols or abstract base classes can describe interchangeable collaborators more explicitly. For now, the important point is the separation of responsibilities and not the fact that each responsibility has become a class.

Neither class needs an inventory, a sales ledger, or an owl.

Putting the shop back together

We still need something to coordinate the sale. First, we give the sales ledger and the notifier small, focused interfaces:

#23

Notice the "pending" state in the ledger. Separating responsibilities does not make the sale process independent of anything else. Inventory has already changed before later steps run. Here, the narrower goal is to make notification failure visible so that the recorded sale can be investigated.

Now we can keep PotionShop as a thin coordinator. It still represents the sale process, but it delegates each part of that process to the object that owns the relevant behaviour:

#24

We can now assemble the shop and sell a potion:

#25

Here’s the output:

🌙 Moon Festival Invisibility 🌙 £10.80
Owl dispatched to Merlin
10.80
{’customer’: ‘Merlin’, ‘potion’: ‘invisibility’, ‘price’: Decimal(’10.80’), ‘notification’: ‘sent’}

If notification fails after the ledger records the sale, the sale remains
visible with a pending notification:

#26

Here’s the output:

Invisibility  £12.00
{’customer’: ‘Merlin’, ‘potion’: ‘invisibility’, 
 ‘price’: Decimal(’12.00’), ‘notification’: ‘pending’}
​

The coordinator still changes if the business process changes. That’s fine: coordinating a sale is its responsibility. But recipe changes belong in RecipeBook, inventory changes belong in Inventory, pricing changes belong in PriceCalculator, label changes belong in LabelMaker, and notification changes belong in Notifier.

The Single Responsibility Principle

This is the idea behind the Single Responsibility Principle:

A class should have one reason to change.

I always find these “official” definitions a bit cryptic when I read them at first. Here, “reason” is about the policy or stakeholder driving the change. The pricing rules may change because the shop owner changes the discounts. The label design may change because the person responsible for packaging wants something different. The notification service may change because the owls are on strike.

That does not mean that every class should have one method. Nor does it mean that every noun in the story needs its own class.

PriceCalculator can have several related methods. Inventory can have several operations. The important question is whether the methods belong to the same coherent responsibility.

The thin PotionShop coordinator is also a class with a responsibility. It coordinates the sale. A change to that business process may quite reasonably affect it.

When unrelated changes keep arriving at the same class, the class becomes harder to test and more likely to break. Extract a collaborator when it may need to change for a different reason from the surrounding code, or when separating it makes an important test simpler. A good class name alone isn’t a reason to create one.

The potion shop did not become troublesome because it had a particular number of lines. It became troublesome because recipes, inventory, pricing, labels, sales, and notifications all had different reasons to change.

A useful question to ask is:

What different kinds of change could affect this class?

If the answers are “the pricing rules”, “the label design”, “the storage system”, and “the notification service”, you may have several responsibilities sharing one home.

And if the project keeps growing, we may need to rethink the floor plan.


This is the first article in a series on the SOLID principles. I’ll look at what each principle means, but more importantly, why it matters and what can go wrong when we ignore it.

Code in this article uses Python 3.14.


Do you want to keep up with how Python programming and the art of making things are changing in this rapidly changing world? Then don’t miss out on the exclusive content for premium members here on The Python Coding Stack.

Subscribe now

I’ll be sharing more about how my programming, my daily work, and everything has changed over the recent months.


For more Python resources, you can also visit Real Python—you may even stumble on one of my own articles or courses there!

Also, are you interested in technical writing? You’d like to make your own writing more narrative, more engaging, more memorable? Have a look at Breaking the Rules.

And you can find out more about me at stephengruppetta.com


Appendix: Code Blocks

Code Block #1
from decimal import Decimal
​
class PotionShop:
    recipes = {
        “invisibility”: {
            “ingredients”: {
                “moonflower”: 2,
                “dragon_scale”: 2,
            },
            “price”: Decimal(”12.00”),
        }
    }
​
    def __init__(self):
        self.stock = {
            “moonflower”: 10,
            “dragon_scale”: 2,
        }
        self.sales = []
​
    def sell(self, potion_name, customer):
        potion = self.recipes[potion_name]
​
        for ingredient, quantity in potion[”ingredients”].items():
            if self.stock[ingredient] < quantity:
                raise ValueError(f”Not enough {ingredient}”)
​
        for ingredient, quantity in potion[”ingredients”].items():
            self.stock[ingredient] -= quantity
​
        price = potion[”price”]
        self.sales.append((customer, potion_name, price))
​
        print(f”{potion_name.title()}\t£{price:.2f}”)
        self.send_owl(
            customer,
            {”potion_name”: potion_name, “price”: price},
        )
​
        return price
​
    def send_owl(self, customer, message_content):
        print(f”Owl dispatched to {customer}”)
        return message_content
Code Block #2
shop = PotionShop()
print(shop.sell(”invisibility”, “Merlin”))
Code Block #3
# ...
​
class PotionShop:
    # ...
​
    def sell(
            self,
            potion_name,
            customer,
            *,
            moon_festival=False,
    ):
        # ...
​
        label = potion_name.title()
        if moon_festival:
            label = f”🌙 Moon Festival {label} 🌙”
​
        print(f”{label}\t£{price:.2f}”)
​
        # ...
    # ...
Code Block #4
shop = PotionShop()
print(shop.sell(”invisibility”, “Merlin”, moon_festival=True))
Code Block #5
# ...
​
class LabelMaker:
    def make(self, potion_name, moon_festival=False):
        label = potion_name.title()
​
        if moon_festival:
            label = f”🌙 Moon Festival {label} 🌙”
​
        return label
​
# ...
Code Block #6
labels = LabelMaker()
assert labels.make(
    “invisibility”,
    moon_festival=True,
) == “🌙 Moon Festival Invisibility 🌙”
Code Block #7
# ...
​
class PotionShop:
    # ...
​
    def __init__(self, label_maker):
        self.label_maker = label_maker
        # ...
​
    def sell(
            self,
            potion_name,
            customer,
            *,
            moon_festival=False,
    ):
        # ...
        label = self.label_maker.make(
            potion_name,
            moon_festival=moon_festival,
        )
        # ...
​
    # ...
Code Block #8
shop = PotionShop(LabelMaker())
Code Block #9
# ...
​
class OwlUnavailable(Exception):
    pass
​
# ...
​
class PotionShop:
    # ...
​
    def __init__(self, label_maker, owl_available=True):
        self.label_maker = label_maker
        self.owl_available = owl_available
        # ...
​
    # ...
​
    def send_owl(self, customer, message_content):
        if not self.owl_available:
            raise OwlUnavailable(
                “The owls are currently unavailable”
            )
        print(f”Owl dispatched to {customer}”)
        return message_content
Code Block #10
shop = PotionShop(LabelMaker())
​
assert shop.sell(”invisibility”, “Merlin”) == Decimal(”12.00”)
Code Block #11
shop = PotionShop(LabelMaker(), owl_available=False)
​
assert shop.sell(”invisibility”, “Merlin”) == Decimal(”12.00”)
Code Block #12
class PotionShop:
    # ...
​
    def sell(
        self,
        potion_name,
        customer,
        *,
        moon_festival=False,
        own_cauldron=False,
    ):
        # ...
​
        price = potion[”price”]
​
        if own_cauldron:
            price = (price * Decimal(”0.90”)).quantize(Decimal(”0.01”))
​
        # ...
​
        return price
​
    # ...
Code Block #13
shop = PotionShop(LabelMaker())
print(
    shop.sell(
        “invisibility”,
        “Merlin”,
        own_cauldron=True,
    )
)
Code Block #14
shop = PotionShop(LabelMaker(), owl_available=False)
print(
    shop.sell(
        “invisibility”,
        “Merlin”,
        own_cauldron=True,
    )
)
Code Block #15
shop = PotionShop(LabelMaker(), owl_available=False)
try:
    print(
        shop.sell(
            “invisibility”,
            “Merlin”,
            own_cauldron=True,
        )
    )
except OwlUnavailable:
    print(”Catching the ‘OwlUnavailable’ Exception”)
    print(shop.stock)
    print(shop.sales)
Code Block #16
class RecipeBook:
    def __init__(self):
        self.recipes = {
            “invisibility”: {
                “ingredients”: {
                    “moonflower”: 2,
                    “dragon_scale”: 2,
                },
                “price”: Decimal(”12.00”),
            }
        }
​
    def find(self, potion_name):
        return self.recipes[potion_name]
Code Block #17
recipe_book = RecipeBook()
assert recipe_book.find(”invisibility”)[”price”] == Decimal(”12.00”)
Code Block #18
class Inventory:
    def __init__(self, stock):
        self.stock = stock
​
    def remove_stock(self, ingredients):
        for ingredient, quantity in ingredients.items():
            if self.stock[ingredient] < quantity:
                raise ValueError(f”Not enough {ingredient}”)
​
        for ingredient, quantity in ingredients.items():
            self.stock[ingredient] -= quantity
​
    def add_stock(self, ingredients):
        for ingredient, quantity in ingredients.items():
            self.stock[ingredient] = self.stock.get(ingredient, 0) + quantity
Code Block #19
inventory = Inventory({
    “moonflower”: 10,
    “dragon_scale”: 2,
})
​
try:
    inventory.remove_stock({
        “moonflower”: 2,
        “dragon_scale”: 3,
    })
except ValueError as error:
    print(error)
​
print(inventory.stock)
Code Block #20
inventory.remove_stock({
    “moonflower”: 2,
    “dragon_scale”: 2,
})
inventory.add_stock({”moonflower”: 5})
print(inventory.stock)
Code Block #21
class PriceCalculator:
    def calculate(self, recipe, own_cauldron=False):
        price = recipe[”price”]
​
        if own_cauldron:
            price = (
                price * Decimal(”0.90”)
            ).quantize(Decimal(”0.01”))
​
        return price
Code Block #22
pricing = PriceCalculator()
print(
    pricing.calculate(
        {”price”: Decimal(”12.00”)},
        own_cauldron=True,
    )
)
Code Block #23
class SalesLedger:
    def __init__(self):
        self.sales = []
​
    def record(self, customer, potion_name, price):
        sale = {
            “customer”: customer,
            “potion”: potion_name,
            “price”: price,
            “notification”: “pending”,
        }
        self.sales.append(sale)
        return sale
​
    def mark_notified(self, sale):
        sale[”notification”] = “sent”
​
class Notifier:
    def send(self, customer, potion_name, price):
        print(f”Owl dispatched to {customer}”)
Code Block #24
class PotionShop:
    def __init__(
        self,
        recipe_book,
        inventory,
        pricing,
        labels,
        ledger,
        notifier,
    ):
        self.recipe_book = recipe_book
        self.inventory = inventory
        self.pricing = pricing
        self.labels = labels
        self.ledger = ledger
        self.notifier = notifier
​
    def sell(self, potion_name, customer, *, moon_festival=False,
             own_cauldron=False):
        recipe = self.recipe_book.find(potion_name)
        self.inventory.remove_stock(recipe[”ingredients”])
​
        price = self.pricing.calculate(
            recipe,
            own_cauldron=own_cauldron,
        )
        label = self.labels.make(
            potion_name,
            moon_festival=moon_festival,
        )
        sale = self.ledger.record(customer, potion_name, price)
​
        print(f”{label}\t£{price:.2f}”)
        self.notifier.send(customer, potion_name, price)
        self.ledger.mark_notified(sale)
​
        return price
Code Block #25
ledger = SalesLedger()
shop = PotionShop(
    RecipeBook(),
    Inventory({”moonflower”: 10, “dragon_scale”: 2}),
    PriceCalculator(),
    LabelMaker(),
    ledger,
    Notifier(),
)
​
print(
    shop.sell(
        “invisibility”,
        “Merlin”,
        moon_festival=True,
        own_cauldron=True,
    )
)
​
assert ledger.sales[0][”notification”] == “sent”
print(ledger.sales[0])
Code Block #26
class FailingNotifier:
    def send(self, customer, potion_name, price):
        raise OwlUnavailable(”The owls are currently unavailable”)
​
ledger = SalesLedger()
shop = PotionShop(
    RecipeBook(),
    Inventory({”moonflower”: 10, “dragon_scale”: 2}),
    PriceCalculator(),
    LabelMaker(),
    ledger,
    FailingNotifier(),
)
​
try:
    shop.sell(”invisibility”, “Merlin”)
except OwlUnavailable:
    pass
​
print(ledger.sales[0])
assert ledger.sales[0][”notification”] == “pending”

stephengruppetta.com

September 28, 2026 05:17 PM UTC


Glyph Lefkowitz

What Would A Serious AI Product Look Like?

One of the issues that I have with the current generation of “AI” products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like.

Even before we get to the tremendous ethical problems with the frontier labs, it is this impression of their composition as a product that makes me feel, constantly, whenever I am interacting with them, that they are less a software product than that they are a grift, a scam designed to make me feel like I am interacting with a product that has capabilities that it simply does not, to try to lull me into a false sense of security that I can trust it.

The frontier labs are of course the worst offenders, but every criticism here applies just as much to Ollama, which (if anything, due to the obviously poorer quality of the available models themselves) needs these features even more than the frontier labs do.

Here, I will set down a few features that might convince me that an LLM-based product, particularly one focused on research or software development, was actually serious about helping me do useful things with it.

Make “Checking For Mistakes” A First-Class Feature

This is the biggest issue, and the major reason that I was inspired to write this post.

It is a truth universally acknowledged, that AIs cannot reliably provide information.

I could cite a ton of news articles and studies about this fact, but there is no need. Every single chatbot admits this, up front, in a fine-print disclaimer as a core part of their user interface. Gemini says “AI can make mistakes, so double-check responses”, Claude says “Claude is AI and can make mistakes. Please double-check responses.1” ChatGPT says “ChatGPT can make mistakes. Check important info.”.

Every time I see that last one, I wonder how I’m supposed to know what “info” is supposed to be “important”.

All of these warnings are all small, gray text, painfully obviously included as legalese to push responsibility back onto the user rather than to help with anything. This is a core limitation of all these products. Checking their output is a part of the workflow for using them that:

  1. you absolutely cannot skip or skimp on without creating risks to yourself and whoever you are conveying its output to, and,
  2. it is very easy to skip or skimp on and you are encouraged at every turn to do so, because “just trust the output” is one of the quickest ways to save time.

A chatbot product that took this weakness seriously, as an actual consideration for using it, would put a checkbox next to every claim in its output. It would be a 2-column worksheet, where you’ve got the LLM output in the first column, and next to it, human notes in the second column, explaining what work went into checking this claim, and a big checkbox that you would only check off after you believe you’d checked its claims thoroughly enough.

Coding assistants would need to have some version of this as well. Right now, this is pushed off into code review, which means it is a dark pattern which subtly encourages the “author”2 to offload this work to their code reviewer without ever looking. Once again, “it’s probably fine, I don’t need to check” is the quickest way to save time and churn out those PRs faster.

It might even be useful for coding harnesses to have some affordance for checking code before it even runs tests. As the vendors themselves have admitted, it’s not just expensive to burn tokens on your “AI”, you also end up burning far more compute on the AI. Being able to check your diffs before sending them over to uselessly exhaust your testing compute cluster would be useful.

If your product tells me that it makes mistakes and I must be the one to check for the mistakes, but then gives me zero tools to check for mistakes, I cannot take it seriously.

More Citations to Check, And More Details

Most chatbots prefer to give an answer, rather than a citation. In my own personal use, I find that when asked to provide a list of citations with clearly marked sources for each one, they will appear to “get bored” halfway through the list and simply stop including citations at some point.

When the bots include citations at all, present them as inline annotations that say nothing but the domain name of the search result, in a font so small that it’s barely legible, and an equally indecipherable icon that is fewer than 16 pixels on a side.

This is backwards.

Now, I am aware that these citations do come from somewhere, and in an attempt to reduce hallucinations, all of the major providers support some form of “grounding”3, and that those little barely-readable citation links are referencing actual structures in the RAG pipeline and not just potentially-hallucinated tokens, but I’m not talking about the underlying machinery in the model, I’m talking about the presentation to the user.

Plus, regardless of whether a snippet of text came from a RAG query, we know that LLMs can never provide an authoritative result; it’s a fundamental limitation of the technology. They can still garble the results of RAG as much as they can misrepresent any other training data. This means that it must never present its results as authoritative.

If you ask an AI to do research queries, every result should be presented as a list of citations. Moreover, the presentation should display each citation as a large object of in its own right, with clearly identified metadata, including not just the site where it was found but its publication date and, if possible, the name of the author. The literal, unmodified quotation (not from RAG, not a summary: a quotation extracted with a regular program and not an LLM) should be front-and-center, larger than any AI-generated text.

If the AI product wants to editorialize or summarize (which should not always be necessary!), the AI-generated text should be presented as small text underneath the citation that has been found, de-emphasized as much as the disclaimer is right now, at the very least until the user has verified that the summary is accurate. Perhaps, for a research project, a “did you read the citation” checkbox might even be helpful.

If your product openly tells me that it will scramble, misrepresent, or omit its citations in its summaries, and I must read the original human-authored citations to be sure, but then gives me no tools to track my reading of those citations or even any way to find them, I cannot take it seriously.

No First-Person Output, No Apologies

There is no reason for a software development or research tool to use first-person language to describe itself. They should not do so. In fact they should not be allowed to do so.

There is also no reason that they should ever apologize. It is a waste of everyone’s time; it’s a waste for the chatbot to generate the apology, it’s a waste for the user to read the apology, and it’s a waste for the user to respond to the apology. Yet they unfailingly do this upon every correction.

The vendors of these tools know that they are routinely causing mental-health crises. In response, they have added non-functional “guard rails” that can still, in 2026, easily be bypassed.4

A product seriously interested in helping with productivity would correct this glaringly obvious flaw, focus on the task at hand, and stop emitting useless verbiage.

In the previous two sections, I tried to focus on ways in which the harness would be constructed differently even if the LLM technology is fundamentally impossible to improve; in this case, I have to assume that the labs have some control over the model itself. But unless they are truly incapable of influencing their output (and all their “benchmarks” and “capabilities” seem to indicate that they can control it very tightly) they ought to be building models that are much less verbose.

More Non-Natural-Language User Interfaces

Although natural language could hypothetically be a powerful interface for interacting with a computer system, the practical upshot of LLM natural language interfaces is that these interfaces are imprecise and repetitive, full of superstitions masquerading as “best practices”. The inputs are a mess and the resulting outputs are a mess.

The general way of addressing this unstructured mess is to allow the chatbot to directly take action in response to the user’s input; in other words to supply it with “tools” via an MCP server. But again, this is backwards. If we cannot even express our intent clearly in the first place, why are we trusting this system to take potentially destructive and harmful actions on our behalf?

Instead, I would expect a product that was seriously invested in helping me accomplish specific tasks, to have user interfaces specific to those tasks. Is it supposed to be able to be a security scanner that can discover OWASP top 10 bugs in a codebase? Have a button for that. Build that functionality into your harness, train it directly into the model, use smaller models that can satisfy that functionality more effectively than throwing it at the planet-sized brain of Fable or whatever.

I’m aware that there are small software startups that do something like this, but they are bolted on to the side of the main model providers’ APIs, not integrated into the core of the product and not using their own models and AI systems to achieve consistent and repeatable results.

Strong Data Provenance Indicators

Chatbots produce data tables pulled from websites, from APIs, from MCP tools or from summarizing and scrambling the user’s input. In order to provide the illusion of a seamless interface, this data is presented in-line regardless of where it comes from. But some of these outputs are produced mechanically via regular old API calls, for example, from the result of calling a tool or querying a website, but presented uniformly.

But there is a huge difference between an authoritative data source being inlined as part of a chatbot conversation, being treated as input by the chatbot, and some ad-hoc hallucinated data being treated as output of the chatbot.

If a product is trying to help me make accurate, empirically-grounded, data-driven decisions, the source of the data is critical.

Integrated into the “check for mistakes” and “verify citations” workflow I described above, there’s a necessary “verify data programmatically” pass as well; to have tools that will treat portions of the output as a regular spreadsheet, allowing regular-old computer arithmetic to verify things and showing where such arithmetic was used, and how.

Better User Control of Reproducibility

Anyone familiar with the technical specifics of LLMs will know that they have a variable called “temperature” which controls the degree of randomness that the LLM uses to produce its outputs. But most users don’t know this, because it isn’t exposed as part of the user interface by default.

This leads to a subjective impression that you asked ChatGPT, and you got ChatGPT’s authoritative answer.

You can’t just set the temperature to zero and still get useful results - I am aware that it does more than just scramble the output at random, and there are perhaps good reasons that simply exposing just a temperature setting would not be that useful to users. But if we followed some more of my earlier recommendations for making more structured UI elements to solve specific problems rather than having long back-and-forth chats where each refinement depends on the previous response, perhaps those elements could also re-play the process so that users can see how reliable the bot is at a particular task and develop a sense of how the stochastic nature of the process actually affects it.

Similarly, if a user is trying to solve the same problem repeatedly with a chatbot, and the chatbot product has numerous computational tools that don’t really have anything to do with the LLM, such as deterministic data-processing tools, then having a way to freeze the non-deterministic parts of the transcript but re-populate a particular data frame with updated information and fork / continue the conversation from there would be a way to avoid introducing pointless additional randomness when you already know what tool you’re trying to use.

The fact that every conversation is presented as this flat chat prompt that doesn’t let me interact with any of the widgets that were previously produced except through more chatting, really makes me feel like the whole product is just doing predatory social-media style “increase time on site” optimization, just trying to lure me into further repetitive and unreliable chats, rather than letting me get in, solve my problem, and get out.

Context Visibility

Managing the LLM context is the ongoing challenge facing organizations that are trying to use “agentic” workflows. Filling up the context with too much information causes well-known problems. In response, advanced LLM users attempting to solve larger problems must break up very long prompts into “skills”, give access to lengthy information via “tools”, and delegating sub-problems to “sub-agents” rather than simply extending a single prompt indefinitely.

All of these strategies have flaws, because even on the largest models, compared to the breadth and depth of knowledge-work problems, LLM contexts are quite small.

And yet, none of these products will show the context to the user by default. There are third-party addons that can show you a simple progress bar but for addressing the premier engineering difficulty with this technology, that is below the bare minimum.

This lack of visibility means that almost all of the tools for extending the context are flying blind. Rather than responding meaningfully to a full context, everyone just kind of guesses how much state they need by guessing and trying over and over again with progressively more elaborate skill and sub-agent layouts. Even managing context compaction ends up being an advanced API-driven workflow5.

A serious product that was trying to help the user understand would not only show “available context” but explain the impact of context compactions, make it easier to see harness-generated prompts, and so on. This would be a first-class feature, combined with the aforementioned reproducibility / replay tools, would allow users to do real experiments to develop an understanding about how to make good use of the context window.

A Sandbox That Actually Works

I’ve been focused on the chatbot interface here because it is the most immediately egregious upon looking at the UI. But the “agentic loop” tools used for coding are equally dangerous, if not more so. Coding tools keep destroying everyone’s data, over the course of years.

These catastrophic incidents that become front-page news are relatively rare compared to the amount of coding-agent use out there. But they also aren’t the only kind of sandbox violation. Coding models will so routinely edit test code instead of the system under test that there are “pro tips” articles all over the web giving you the flawed advice to simply ask the agent not to cheat. News write-ups of the catastrophic incidents themselves will also offer glib and wrong advice, like “use a docker container”. That might prevent it from literally deleting your operating system, but it won’t prevent it from destroying all the local work you have in your codebase (it needs access to a checkout, after all!)

There is a flurry of activity in the infosec space where people are rushing to plug the gaps left by these coding harnesses. Everyone’s got their own version of an MCP approval gateway where you can optionally place a proxy between your agent and your production infrastructure.

In the best case, though, all these mitigations and proxies and prompts simply turn the user into an auto-approval automaton, hitting Y, Y, Y, Y over and over again, until you finally are driven mad and hit “yes to all”, turn on full-auto mode and submit yourself to the void. With nothing between your personal vigilance and disaster, there are no workflows left beyond decrementing your own vigilance until there’s nothing left and then hoping the disaster never arrives.

The fact that some mitigations exist that can be deployed by extra-cautious users does not change the fact that “agentic coding” is an unsafe-by-default technology deployed without concern or guidance. Every frontier lab has tied a spring-loaded shotgun to a dog; the fact that dog owners can publish thoughtful blog posts explaining how you can teach your dog the basics of gun safety or how you can have your dogs play in a bullet-proof room does not mitigate the fact that the product should not have been allowed in the first place, nor should it continue to exist without VERY strong security controls.

I might believe that a frontier lab were seriously interested in providing developers with a useful tool if they shipped something that had safety built-in.

That means tools in the harness, detached from any LLM, independent of the prompt, that could:

In the same way that I suggested above that research-based tasks should have a way of re-issuing prompts to determine how reproducible a result is, or whether other sources might be found, agent-based tasks should have a way of being executed against mock services for popular APIs, so that the verification can match both on the front-end (review the plan for making the API calls before they’re executed) and the back end (review the API calls that were issued to the mock service and verify that they matched).

Instead, the frontier labs provide us products that are disasters out of the box, give us “best practices” to build massive and elaborate, as well as incomplete and error-prone, security perimeters of our own design. Then they blame “operator error” when it inevitably goes wrong. I cannot believe that these design choices are intended to help us be productive.

Bonus: Human Processes

Organizations deploying AI also frequently come across as unserious, for similar reasons. In 2023, naive exuberance could perhaps be forgiven. But today, as we near the close of 2026, there are several well-known problems, that have been extremely well-covered in the press. None of these things should be surprising, but most orgs deploying these tools are still just letting them rip and hoping it all works out.

Organizations deploying these tools would need at least three kinds of major modifications to their internal processes, if they wanted to be serious about using them safely:

1. Shift Rotations to Prevent Vigilance Decrement

There have been several high-profile incidents where software developers’ gradual acquiescence to accepting LLM output have lead to serious economic consequences for the companies deploying them, perhaps best typified by Amazon’s “millions of lost orders” due to a gradual decay of their engineering processes from LLM use.

These outages, and other AI-related failures, are due to the difficulty of maintaining focus on the same problems. In other words, as I described above, vigilance decrement is a constant problem, because AI outputs are most often correct, but continue to be incorrect in surprising and non-intuitive ways. As I have previously written, you cannot trust yourself to catch every bug with code review, and LLM output.

Aviation, for example, has very strict rules around rest requirements. There is also a specific rule that “No certificate holder may operate an aircraft without a second in command if that aircraft has a passenger seating configuration, excluding any pilot seat, of ten seats or more.”. Other safety-critical professions have similar rules.

And yet, even in the age of the supposed “AI revolution”, most software teams are still assigning every engineer a full feature load, not planning for any rest, and telling people to review code whenever they happen to have some “free time”.

Maintenance of vigilance has to be your top priority. Regular, scheduled, inviolable rest periods where people do work without AI assistance, and are not exposed to any AI output for review or otherwise, would be crucial in order to stay mentally sharp enough.

The tools themselves should have this sort of thing built in. The mistake-review process described above should have a periodic spot-check mode where a second reviewer periodically reviews a chatbot log, doing their own independent verification of claims, to see if they spot the same errors. This could provide a feedback loop to determine how much rest is necessary to maintain continuous attention and actually spot hallucinations.

2. Skill Practice To Prevent Skill Loss

It is also well-known that AI use leads to AI reliance, and AI reliance leads to skill loss.

I like to use the analogy to dockworkers at a seaport6 adopting automation.

If you employ dockworkers to load and unload ships all day long, they are going to be getting tons of exercise. They will be able to lift heavy objects on demand, whenever. They might have plenty of health problems and injuries from this type of work, but “lack of exercise” will not be a problem.

With the development of standardized container ships and mechanized cranes, you are going to be changing their job description substantially: now they mostly spend all day sitting in a small cubicle moving a control lever back and forth, not lifting heavy stuff. They will get worse at lifting heavy objects.

In this analogy however, the cranes are not all that reliable. We know they break, and they drop their payloads sometimes, and the stuff needs to be manually moved. But this only happens a few times a week, at most. If you need whoever is driving the crane to be able to jump out at any moment and still move stuff around manually, then you need to make an affordance for that. You need to give them time to go to the gym and do some lifting for practice, or every crane failure is going to be a major emergency.

An organization doing an AI transformation would also need a massive increase to learning & development budget, both in terms of resources and in terms of schedule. If your people are going to lose skills because they’ve lost regular practice in the incidental course of doing their duties, then they are going to need deliberate, intentional, non-incidental practice of those skills to keep them sharp.

But rather than trying to accommodate new workflows and give time for people to adjust, most AI mandates are simply dropped on workers like a ton of bricks, with no time to adapt and no affordance for maintaining their skills. Operate the crane and stay fit and healthy and ready to switch back to manual lifting at any time and then get back in the crane cockpit right afterwards. Don’t mess up.

Then an accident happens and everyone is surprised, as if this process weren’t practically designed to produce a terrible result.

3. Mental Health Resources to Deal with Mental Health Risks

AI psychosis often begins with practical problem-solving, and beyond that, it can start specifically at work. Not to mention the more pedestrian condition of “AI brain fry”.

If you are mandating your employees to use a hazardous tool that may seriously and directly damage their mental health, you need trainings and resources. You need in-house therapists and you need to be making sure to check in with people actively to make sure that this is not happening.

Again, the tool itself ought to have some way of dealing with this. An occasional “take a break” popup is easily dismissed; they need a user-visible AI personal dosimeter so you can see your cumulative usage over time.

I don’t even know if “usage over time” is a sufficient metric to gauge risk. Maybe if your work chatbot start to talk about resonance too much, unless you literally work as an acoustic engineer, that should be flagged for someone.

We are, again, years into dealing with these tools, and we know these risks exist. Yet no serious mitigations are provided. Not even any way of measuring the risk exposure.

And More

There are also many other risks associated with the technology. There are intellectual property risks with the foundation models, due to recklessness with their training data. There are existential financial risks associated with the infrastructure build-out. The extent to which most “open” models are simply derivatives of frontier models is an open question.

What I Think

If any one of these things were regularly overlooked by AI vendors or users, that would be a totally normal product oversight. Room for improvement for the next version, but nothing catastrophic.

Shipping without any of them doesn’t seem like lean product management, it seems like a careless attitude towards risk and a product design philosophy oriented entirely towards short-term demos, with no regard for how to realize actual productivity gains.

Furthermore, being available for years without anything like these features, despite hundreds of incidents demonstrating the risks, with hundreds of billions of dollars of funding, makes it seem to me like if they were to add all the features that would make their product actually safe and hypothetically useful, these features would reveal that it is actually not an improvement to productivity.

In the year since I first wrote about measuring the cost/benefit ratio of AI, I have heard from numerous people who have shown this to management to try to illustrate why their AI initiatives — like almost all AI initiatives — were either failing or burning out their engineers.

I’ve also heard from lots of people that have told me that it’s obviously useful and they don’t need to measure so carefully, because they are getting lots of work done that they couldn’t have otherwise.7

I have yet to hear from a single person who has said “yeah, we measured according to your methodology8, and it turns out that our AI work is going great and that our ratio is 0.75”.

Obviously, I cannot say for sure why this is; absence of evidence is not evidence of absence. But at this point I think the null hypothesis is that AI tools provide, in aggregate, zero value. They make mistakes too often, and the externalities they produce are so bad and so difficult to control that even before we get to the places where they are just physically poisoning people, even the negative effects on their direct users end up cancelling out whatever benefit to they provide to their organizations.

If I were wrong, then including tools to measure an AI’s effectiveness at the tasks their users are actually trying to accomplish, rather than meaningless benchmarks, would show big productivity gains. The frontier labs would be champing at the bit to add such features, and crowing about their fantastic results.

I think the labs know that if they did that, it would present a grim picture to their users. Such tools would let their users see that it’s making mistakes much more often than they realized, that they’re spending much more time with it than they want to be, and that it’s just generally not fit for purpose.

If they prove me wrong by adding in all of these safety mechanisms, and in the process, they make all of their AI technology less harmful, I’ll be thrilled to be debunked.


Acknowledgments

Thank you to my patrons who are supporting my writing on this blog. If you like what you’ve read here and you’d like to read more of it, or you’d like to support my various open-source endeavors, you can support my work as a sponsor!


  1. It is also interesting that for the next section, sometimes it seems that Claude’s disclaimer is “Please double-check cited sources.” instead. ↩

  2. ... by which I mean the “prompter”, since authorship is not what’s happening here. ↩

  3. Claude has the “citations API”, Google has various different kinds of “grounding” against its own APIs, and I guess Microsoft can check OpenAI’s homework if you want. ↩

  4. Given the relatively slow speed of the justice system and the mainstream press around the world, we probably will not hear about whether people are managing to incidentally break through these guard rails to self harm right now, but there are no shortage of stories still being reported right now where people were still doing just that, such as in this story where the effect of the much vaunted “guard rails” in 2025 was that if you wanted it to write you a suicide note, it would refuse twice but acquiesce on the third try. I don’t see any reason to believe this fundamental issue has been addressed in the meanwhile, since it had been happening for years at that point. ↩

  5. In this tutorial we can also see an incredibly rosy scenario presented, where a long-running workflow effortlessly compresses all of the necessary information into the new context, even if it uses a lower-fidelity model to do so, rather than the tangled and gnarly problem of problems which really are too big to fit in the context, which is to say, “most real-world problems”. This presents the context limit instead as a minor speedbump to be worked around rather than the fundamental flaw in LLM tooling. ↩

  6. A heavily fictionalized seaport. This is not how actual dockworkers work. This is a simplistic metaphor about incidental benefits of instrumental tasks, it is not supposed to delve deeply into the mechanics of maritime shipping. In particular I know that cranes are more reliable than this and this is not actually how you would respond to a crane malfunction anyway. Feel free to share fun facts about maritime shipping if that is your special interest but please do not @ me to correct this metaphor. ↩

  7. To my knowledge, none of their publicly-traded employers have posted a measurable improvement to efficiency outside the margin of error. ↩

  8. Or any similar methodology. I don’t need people to adopt the exact practice that I proposed there. ↩

September 28, 2026 04:27 AM UTC

September 27, 2026


LernerPython blog, from Reuven Lerner

Pandas dropna with thresh: Drop only rows with too many NaN values

Pandas dropna with thresh: Drop only rows with too many NaN values

You can remove NaN from a Python Pandas data frame with dropna, but be careful: It removes rows with even one NaN, which can be overkill.

Pass “thresh” to allow some rows to remain, even if they contain NaN:

df.dropna(thresh=2)  # keeps rows with 2 or more non-NaN values

The post Pandas dropna with thresh: Drop only rows with too many NaN values appeared first on LernerPython.

September 27, 2026 06:00 AM UTC

September 26, 2026


Python Insider

The Python documentation is now available in German

You can now read the Python documentation online in German!

September 26, 2026 12:00 AM UTC

September 25, 2026


Rodrigo Girão Serrão

TIL #146 – Using maturin through uv

Today I learned how to setup a Rust project that can be called from Python with PyO3 and maturin through uv.

When you follow the PyO3 getting started guide to create a simple Rust project that can be called from Python, the instructions you get assume you'll use a global Python installation to create a virtual environment and to install maturin into it. You can use uv through maturin, but if your project also has a Rust binary, things may break.

When you run a command like cargo run, cargo will see the dependency on PyO3 and it will then look for a Python installation. If you have no global Python installations — because you do everything through uv — or if your global installations aren't setup exactly like a vanilla, default installation, PyO3 might fail.

The fix is simple. In .cargo/config.toml add the environment variable PYO3_PYTHON that points to the Python inside your virtual environment:

# .cargo/config.toml
[env]
PYO3_PYTHON = { value = ".venv/bin/python", relative = true }

How to set up a Rust + Python project with PyO3 and maturin through uv

Here are all the steps to set up a Rust project that can be compiled into a binary executable and that can also be used from within Python:

% cargo new calculator
% cd calculator

Create the file lib.rs:

// lib.rs
pub fn add(a: i32, b: i32) -> i32 {
    a + b
}

#[pyo3::pymodule]
mod calculator {
    use pyo3::prelude::*;

    #[pyfunction]
    fn add(a: i32, b: i32) -> PyResult<i32> {
        Ok(crate::add(a, b))
    }
}

And update the file main.rs to depend on your calculator:

use calculator::add;

fn main() {
    println!("{}", add(1, 2));
}

If you run cargo run, you should get the result 3:

% cargo run
3

Add the PyO3 dependency from the Rust side:

% cargo add pyo3 -F abi3-py38

Update Cargo.toml to configure your crate type so it can be compiled for the Rust binary and for the Python bridge:

# Cargo.toml
# ...

[lib]
name = "calculator"
crate-type = ["cdylib", "rlib"]

Now, create a minimal pyproject.toml:

# pyproject.toml

[project]
name = "calculator"
version = "0.1.0"

[build-system]
requires = ["maturin>=1.0,<2.0"]
build-backend = "maturin"

Add the dependency on maturin and run it:

% uv add maturin
% uv run maturin develop
# ...
✏️ Setting installed package as editable
🛠 Installed calculator-0.1.0

Run Python with uv run python and test your package:

>>> from calculator import add
>>> add(3, 4)
7

At this point you're happy that you can use your Rust code from Python and may not realise that cargo run may no longer work, complaining about Python frameworks, not finding whatever it needs to link, or other weird errors.

Configure cargo to use the Python installation from the virtual environment by adding the file .cargo/config.toml:

# .cargo/config.toml
[env]
PYO3_PYTHON = { value = ".venv/bin/python", relative = true }

Try running cargo run again and note that everything still works:

% cargo run
3

Fixing issues with PyO3 +

...

September 25, 2026 08:34 AM UTC


LernerPython blog, from Reuven Lerner

Pandas dropna: The best way to remove NaN values from a series

Pandas dropna: The best way to remove NaN values from a series

Coming to Python Pandas from NumPy? You’ll reach for np.isnan:

s = Series([10, np.nan, 30])
s.loc[ ~np.isnan(s) ] # returns 10, 30

Unfortunately, this works. Better, use s.isna (or s.isnull).

But the best way to drop NaN? Use dropna:

s.dropna()  # returns 10, 30

The post Pandas dropna: The best way to remove NaN values from a series appeared first on LernerPython.

September 25, 2026 06:00 AM UTC